Spectral Scaling Laws of Muon: Why Your Optimizer Accuracy Must Scale with Your Model
Spectral Scaling Laws of Muon
This paper introduces the first systematic study of the singular value spectrum of the Muon optimizer's momentum buffers across varying model scales (77M to 2.8B parameters). It uncovers "Spectral Scaling Laws," demonstrating that momentum singular values follow predictable power laws with model size, enabling layer-aware optimization of the Newton-Schulz (NS) iteration.
TL;DR
The Muon optimizer—now the engine behind SOTA models like DeepSeek-V4 and GLM-5—replaces standard Adam updates with orthonormalized directions. However, its core approximation tool, the Newton-Schulz (NS) iteration, has a "dark corner": it ignores directions with small singular values. This paper reveals that as models grow, these singular values shrink according to precise power laws. The takeaway? Late layers in frontier models will fail to train properly unless we increase optimizer iterations or tune them layer-by-layer.
Problem & Motivation: The Orthonormalization Gap
Standard LLM training is dominated by AdamW, but Muon has proven to be up to 2x more compute-efficient. Muon works by ensuring update matrices are orthonormal. Because exact orthonormalization (via SVD) is too slow, Muon uses the Newton-Schulz (NS) iteration, a polynomial approximation.
The catch is that NS is an approximation of the sign function. If a singular value is too close to zero, NS fails to push it to 1, effectively "killing" that direction's contribution to learning. Until now, researchers didn't know if this was a theoretical ghost or a pending disaster at scale.
Methodology: Tracking the Momentum Spectrum
To solve this, the authors tracked the singular value quantiles ( to ) of the momentum buffers in GPT-2 models from 77M to 2.8B parameters.
1. The Stabilization Phenomenon
They discovered that the momentum spectrum isn't chaotic. After a short "burn-in" period, the singular value distribution stabilizes and remains constant for the rest of training.
Figure: After a brief transient phase, singular value quantiles stabilize at values that decrease predictably as model size grows.
2. The Spectral Scaling Laws
By plotting these stabilized values against model size on a log-log scale, the authors found a remarkably clean linear relationship. Every layer follows a power law:
Crucially, the exponent is layer-dependent.
- Mid-layers: Scale mildly ().
- Final layers: Scale aggressively ().
Figure: The "Spectral Scaling Laws" showing different decay rates for different depths in the network.
Experiments: How Much Accuracy Do We Actually Need?
The authors performed "rank-p" ablation studies to see how many directions actually matter. If you only orthonormalize the top 10% of directions, performance tanks. If you hit the top 50%, you get nearly full Muon performance.
Using this "50% rule," they extrapolated their laws to a hypothetical 300B parameter model:
- For a mid-layer, a standard 5-step NS (common in academic code) is still plenty.
- For the final MLP layer, the singular values shrink so much that 5-step NS fails completely. You would need a 10-step NS (like that used in DeepSeek-V4) to keep training stable.
Figure: The Newton-Schulz map. Values below the "cliff" remain unorthonormalized, leading to update suppression.
Depth Insights & Conclusion
This work transforms optimizer tuning from a "trial-and-error" headache into a principled engineering task.
Core Takeaways:
- Layer-Awareness is Key: Applying 10 NS iterations to every layer is a waste of FLOPs. Most layers only need 5; only the tail-end of the model needs the heavy lifting.
- Predictability of LLMs: This is yet another "law" in the LLM universe, proving that even the internal linear algebra of our optimizers follows predictable scaling paths.
Limitations: The study focused on GPT-2/Dense architectures. Future work is needed to see if Mixture-of-Experts (MoE) or newer architectures like Mamba follow the same spectral decay patterns.
