Nemotron 3 Ultra: Hardware-Software Co-Design for the Era of 1M-Token Autonomous Agents

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning NVIDIA

Summary
Problem
Method
Results
Takeaways
Abstract

Nemotron 3 Ultra is a 550B parameter Mixture-of-Experts (MoE) model (with 55B active parameters) utilizing a hybrid Mamba-Attention architecture. Pre-trained on 20 trillion tokens and optimized for agentic reasoning, it achieves state-of-the-art accuracy with up to 6x higher inference throughput compared to similar-scale Transformers like Qwen-3.5 and GLM-5.

TL;DR

NVIDIA's Nemotron 3 Ultra is a 550B parameter behemoth that signals the end of the "Attention-only" dominance for large-scale reasoning. By blending Mamba-2 (State Space Models) with Transformer Attention and a LatentMoE architecture, it achieves a 6x throughput leap over rivals like GLM-5 and Qwen-3.5. It is purpose-built for "Agentic Reasoning"—tasks requiring long-horizon code generation, tool use, and 1M-token context retrieval.

The Architectural Shift: Why Mamba-Attention?

The primary bottleneck for autonomous agents isn't just "intelligence"; it's the KV Cache. In pure Transformers, a 1M-token context consumes hundreds of gigabytes of HBM, leaving no room for batching. Nemotron 3 Ultra solves this through a Hybrid Mamba-Attention stack:

  • Mamba-2 Layers: Handle the bulk of sequence modeling with constant-time decoding complexity.
  • Sparse Attention Anchors: Periodically "reset" the context to maintain global dependency accuracy that pure SSMs sometimes lose.
  • LatentMoE: Instead of standard granular MoE, LatentMoE trades hidden-dimension width for more routed experts, maximizing accuracy per FLOP.

Model Architecture Figure 2: The layer pattern of Nemotron 3 Ultra, alternating Mamba-2 and Attention blocks scaled by LatentMoE.

Methodology: Training at the Limit

Training a 550B model is notoriously unstable. NVIDIA shares rare "war stories" of training divergences at 8T and 16T tokens.

  1. NVFP4 Pre-training: To fit training into hardware, they used 4-bit floating point precision—the largest stable demonstration to date.
  2. MTP Boosting: Multi-token prediction (MTP) was used not just as a loss head, but to train the model to "speculate" ahead, improving inference speed by nearly 3x.
  3. Post-training (MOPD): Unlike standard RLHF, NVIDIA used Multi-teacher On-Policy Distillation. They trained "specialist" teachers for coding, math, and web-search, then distilled their collective wisdom into the Ultra student.

Performance: Breaking the Throughput-Accuracy Frontier

In agentic benchmarks, throughput is the currency of success. Because agents must "think" through many loops, slow models are unusable.

Throughput Comparison Figure 1: Nemotron 3 Ultra maintains SOTA accuracy while obliterating the throughput of open weights models.

Key Benchmarks:

  • IOI 2025: 570.0 (Gold Medal Human Level).
  • SWE-Bench Verified: 71.9% (Surpassing many 1T+ parameter models).
  • RULER (1M Context): 94.7% accuracy at full 1M token length.

Quantization and Inference Infrastructure

Deployment was optimized for the GB200 NVL72 rack. By using NVFP4 (4-bit) for routed experts and FP8 for shared experts, the model fits into a single NVLink domain. This hardware-software affinity allows for:

  • Topology-Aware Placement: Ensuring experts stay within high-speed NVLink domains to avoid InfiniBand latency.
  • SSM Cache Optimization: Quantizing the Mamba state to 16-bit with stochastic rounding to prevent precision loss over long sequences.

Critical Insight: The "Recovery" Rate

One of the paper's most fascinating sections is the MOPD Recovery Rate. It shows that while a student model (RLVR) might start far behind a specialized teacher (e.g., an "Office Worker" expert), the MOPD process recovers over 86% of that gap. This proves that an "All-rounder" model can indeed ingest specialized expert behaviors without collapsing other capabilities.

Conclusion

Nemotron 3 Ultra isn't just a bigger model; it's a blueprint for the next generation of LLMs. By offloading sequence history to Mamba and specializing through MOPD, NVIDIA has balanced the trade-off between the "memory" of Transformers and the "speed" of SSMs. For developers building autonomous agents, this is likely the new gold standard for open-weight efficiency.

Find Similar Papers

Try Our Examples

  • Find recent papers investigating the training stability of hybrid Mamba-Transformer models at the scale of 500B+ parameters.
  • What are the underlying theoretical foundations of LatentMoE, and how does it specifically reduce the inference bottleneck of traditional Mixture-of-Experts?
  • Examine how Multi-teacher On-Policy Distillation (MOPD) compares to Proximal Policy Optimization (PPO) in aligning LLMs for complex, multi-turn agentic workflows.
Contents
Nemotron 3 Ultra: Hardware-Software Co-Design for the Era of 1M-Token Autonomous Agents
1. TL;DR
2. The Architectural Shift: Why Mamba-Attention?
3. Methodology: Training at the Limit
4. Performance: Breaking the Throughput-Accuracy Frontier
5. Quantization and Inference Infrastructure
6. Critical Insight: The "Recovery" Rate
7. Conclusion