Nemotron 3 Ultra: Hardware-Software Co-Design for the Era of 1M-Token Autonomous Agents
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning NVIDIA
Nemotron 3 Ultra is a 550B parameter Mixture-of-Experts (MoE) model (with 55B active parameters) utilizing a hybrid Mamba-Attention architecture. Pre-trained on 20 trillion tokens and optimized for agentic reasoning, it achieves state-of-the-art accuracy with up to 6x higher inference throughput compared to similar-scale Transformers like Qwen-3.5 and GLM-5.
TL;DR
NVIDIA's Nemotron 3 Ultra is a 550B parameter behemoth that signals the end of the "Attention-only" dominance for large-scale reasoning. By blending Mamba-2 (State Space Models) with Transformer Attention and a LatentMoE architecture, it achieves a 6x throughput leap over rivals like GLM-5 and Qwen-3.5. It is purpose-built for "Agentic Reasoning"—tasks requiring long-horizon code generation, tool use, and 1M-token context retrieval.
The Architectural Shift: Why Mamba-Attention?
The primary bottleneck for autonomous agents isn't just "intelligence"; it's the KV Cache. In pure Transformers, a 1M-token context consumes hundreds of gigabytes of HBM, leaving no room for batching. Nemotron 3 Ultra solves this through a Hybrid Mamba-Attention stack:
- Mamba-2 Layers: Handle the bulk of sequence modeling with constant-time decoding complexity.
- Sparse Attention Anchors: Periodically "reset" the context to maintain global dependency accuracy that pure SSMs sometimes lose.
- LatentMoE: Instead of standard granular MoE, LatentMoE trades hidden-dimension width for more routed experts, maximizing accuracy per FLOP.
Figure 2: The layer pattern of Nemotron 3 Ultra, alternating Mamba-2 and Attention blocks scaled by LatentMoE.
Methodology: Training at the Limit
Training a 550B model is notoriously unstable. NVIDIA shares rare "war stories" of training divergences at 8T and 16T tokens.
- NVFP4 Pre-training: To fit training into hardware, they used 4-bit floating point precision—the largest stable demonstration to date.
- MTP Boosting: Multi-token prediction (MTP) was used not just as a loss head, but to train the model to "speculate" ahead, improving inference speed by nearly 3x.
- Post-training (MOPD): Unlike standard RLHF, NVIDIA used Multi-teacher On-Policy Distillation. They trained "specialist" teachers for coding, math, and web-search, then distilled their collective wisdom into the Ultra student.
Performance: Breaking the Throughput-Accuracy Frontier
In agentic benchmarks, throughput is the currency of success. Because agents must "think" through many loops, slow models are unusable.
Figure 1: Nemotron 3 Ultra maintains SOTA accuracy while obliterating the throughput of open weights models.
Key Benchmarks:
- IOI 2025: 570.0 (Gold Medal Human Level).
- SWE-Bench Verified: 71.9% (Surpassing many 1T+ parameter models).
- RULER (1M Context): 94.7% accuracy at full 1M token length.
Quantization and Inference Infrastructure
Deployment was optimized for the GB200 NVL72 rack. By using NVFP4 (4-bit) for routed experts and FP8 for shared experts, the model fits into a single NVLink domain. This hardware-software affinity allows for:
- Topology-Aware Placement: Ensuring experts stay within high-speed NVLink domains to avoid InfiniBand latency.
- SSM Cache Optimization: Quantizing the Mamba state to 16-bit with stochastic rounding to prevent precision loss over long sequences.
Critical Insight: The "Recovery" Rate
One of the paper's most fascinating sections is the MOPD Recovery Rate. It shows that while a student model (RLVR) might start far behind a specialized teacher (e.g., an "Office Worker" expert), the MOPD process recovers over 86% of that gap. This proves that an "All-rounder" model can indeed ingest specialized expert behaviors without collapsing other capabilities.
Conclusion
Nemotron 3 Ultra isn't just a bigger model; it's a blueprint for the next generation of LLMs. By offloading sequence history to Mamba and specializing through MOPD, NVIDIA has balanced the trade-off between the "memory" of Transformers and the "speed" of SSMs. For developers building autonomous agents, this is likely the new gold standard for open-weight efficiency.
