MoonMath.ai builds the performance layer for Physical AI.
We are a small team of mathematicians and engineers building production-grade acceleration for the next wave of AI systems via low-level algorithms and systems engineering.
Zro is MoonMath's private inference platform for coding agents. It gives developers fast access to open-weight models through a single endpoint, built for long-context, multi-turn workflows.
- Private by default: Zero request retention and no training on customer data.
- Built for agents: Connect Claude Code, Codex, OpenCode, Cline, and more in minutes.
- Open models, one endpoint: Use capable open-weight coding models without operating the serving stack yourself.
- Performance engineered: Powered by MoonMath's compression, custom kernels, and hardware-aware deployment.
MoonLite is our toolkit for accelerating large generative models:
- LiteAttention: temporal sparse attention for video diffusion.
- LiteLinear: decomposed modules that replace standard FFN layers.
- BackLite: FlashAttention 3-based backward-pass acceleration using sparse gradient approximation.
- LiteRunner: experiment runner for generative models, with local and Weights & Biases tracking.