Spectral Scaling Laws of Muon: Why Your Optimizer Accuracy Must Scale with Your Model

Spectral Scaling Laws of Muon

2026-06-01
Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces the first systematic study of the singular value spectrum of the Muon optimizer's momentum buffers across varying model scales (77M to 2.8B parameters). It uncovers "Spectral Scaling Laws," demonstrating that momentum singular values follow predictable power laws with model size, enabling layer-aware optimization of the Newton-Schulz (NS) iteration.

TL;DR

The Muon optimizer—now the engine behind SOTA models like DeepSeek-V4 and GLM-5—replaces standard Adam updates with orthonormalized directions. However, its core approximation tool, the Newton-Schulz (NS) iteration, has a "dark corner": it ignores directions with small singular values. This paper reveals that as models grow, these singular values shrink according to precise power laws. The takeaway? Late layers in frontier models will fail to train properly unless we increase optimizer iterations or tune them layer-by-layer.

Problem & Motivation: The Orthonormalization Gap

Standard LLM training is dominated by AdamW, but Muon has proven to be up to 2x more compute-efficient. Muon works by ensuring update matrices are orthonormal. Because exact orthonormalization (via SVD) is too slow, Muon uses the Newton-Schulz (NS) iteration, a polynomial approximation.

The catch is that NS is an approximation of the sign function. If a singular value is too close to zero, NS fails to push it to 1, effectively "killing" that direction's contribution to learning. Until now, researchers didn't know if this was a theoretical ghost or a pending disaster at scale.

Methodology: Tracking the Momentum Spectrum

To solve this, the authors tracked the singular value quantiles ( to ) of the momentum buffers in GPT-2 models from 77M to 2.8B parameters.

1. The Stabilization Phenomenon

They discovered that the momentum spectrum isn't chaotic. After a short "burn-in" period, the singular value distribution stabilizes and remains constant for the rest of training.

Model Architecture and Quantile Dynamics Figure: After a brief transient phase, singular value quantiles stabilize at values that decrease predictably as model size grows.

2. The Spectral Scaling Laws

By plotting these stabilized values against model size on a log-log scale, the authors found a remarkably clean linear relationship. Every layer follows a power law:

Crucially, the exponent is layer-dependent.

  • Mid-layers: Scale mildly ().
  • Final layers: Scale aggressively ().

Scaling Law Power Fit Figure: The "Spectral Scaling Laws" showing different decay rates for different depths in the network.

Experiments: How Much Accuracy Do We Actually Need?

The authors performed "rank-p" ablation studies to see how many directions actually matter. If you only orthonormalize the top 10% of directions, performance tanks. If you hit the top 50%, you get nearly full Muon performance.

Using this "50% rule," they extrapolated their laws to a hypothetical 300B parameter model:

  • For a mid-layer, a standard 5-step NS (common in academic code) is still plenty.
  • For the final MLP layer, the singular values shrink so much that 5-step NS fails completely. You would need a 10-step NS (like that used in DeepSeek-V4) to keep training stable.

NS Failure Regime Figure: The Newton-Schulz map. Values below the "cliff" remain unorthonormalized, leading to update suppression.

Depth Insights & Conclusion

This work transforms optimizer tuning from a "trial-and-error" headache into a principled engineering task.

Core Takeaways:

  • Layer-Awareness is Key: Applying 10 NS iterations to every layer is a waste of FLOPs. Most layers only need 5; only the tail-end of the model needs the heavy lifting.
  • Predictability of LLMs: This is yet another "law" in the LLM universe, proving that even the internal linear algebra of our optimizers follows predictable scaling paths.

Limitations: The study focused on GPT-2/Dense architectures. Future work is needed to see if Mixture-of-Experts (MoE) or newer architectures like Mamba follow the same spectral decay patterns.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply layer-wise adaptive compute or precision to deep learning optimizers like Shampoo or SOAP.
  • Who first proposed the Muon optimizer for neural network training and what was the theoretical justification for orthonormalized updates?
  • Are there studies investigating if Spectral Scaling Laws hold true for Mixture-of-Experts (MoE) architectures or Vision Transformers?
Contents
Spectral Scaling Laws of Muon: Why Your Optimizer Accuracy Must Scale with Your Model
1. TL;DR
2. Problem & Motivation: The Orthonormalization Gap
3. Methodology: Tracking the Momentum Spectrum
3.1. 1. The Stabilization Phenomenon
3.2. 2. The Spectral Scaling Laws
4. Experiments: How Much Accuracy Do We Actually Need?
5. Depth Insights & Conclusion