SF-NorMuon: Escaping the Learning Rate Schedule Trap with Spectral Optimization
Anytime Training with Schedule-Free Spectral Optimization
The paper introduces SF-NorMuon, a schedule-free spectral optimizer designed for anytime neural network training. It combines spectral optimization (using the polar decomposition of the gradient) with a schedule-free framework to achieve SOTA performance on 125M and 772M language models, matching per-horizon tuned AdamW baselines.
TL;DR
Training large language models (LLMs) usually feels like a gamble: you pick a training horizon (e.g., 2 Trillion tokens), set a cosine decay schedule, and hope your data estimates were right. If you stop early or want to keep going, your model is sub-optimal. SF-NorMuon changes this by removing the schedule entirely. By combining Spectral Optimization (learning from the gradient's singular structure) with a Schedule-Free framework, it matches the performance of perfectly tuned AdamW benchmarks at any point in time, while solving the long-term stability issues that plagued previous anytime optimizers.
The Horizon Problem: Why Schedules Hold Us Back
Most modern LLMs use Cosine Learning Rate Decay. The learning rate starts high and hits near-zero exactly at the end of the planned training.
The Pain Point:
- Path Dependence: If you plan for 100B tokens but stop at 50B, your model performs worse than a model specifically scheduled for 50B.
- Re-tuning Cost: Changing the budget means re-tuning the schedule.
- SF-AdamW's Gap: Previous attempts at "Schedule-Free" training (like SF-AdamW) existed but often lagged behind the "gold standard" of a well-tuned scheduled run.
Figure 1: SF-NorMuon (green) matches the "stars" (tuned AdamW) across all horizons, whereas SF-AdamW (blue) consistently underperforms.
Methodology: Spectral Geometry meets Anytime Stability
1. Why Spectral? (The Muon Insight)
Standard AdamW treats weight matrices as flat vectors. This ignores their physical role as linear operators. SF-NorMuon uses the Spectral Norm geometry. Instead of entry-wise scaling, it computes the Polar Decomposition of the gradient:
abla f(W))$$ By forcing all singular values of the update to be 1, the optimizer makes uniform progress across all directions in the weight space, regardless of the Hessian's flatness. ### 2. The Weight Decay "Gotcha" The authors discovered a critical flaw in prior schedule-free theory: in non-convex, long-horizon training, the "fast" sequence ($z_t$) eventually diverges if weight decay is applied at the interpolation point ($y_t$). **The Solution**: Apply weight decay directly to the fast iterate $z_t$. This keeps the weights bounded and allows the model to stay in a "Quasi-Steady State" for tens of billions of tokens.  *Figure 2: Without decay at Z (green), anytime methods eventually diverge (blue/orange).* ### 3. The SF-NorMuon Algorithm The algorithm maintains three sequences: - **Z (Fast iterate)**: Moves at a constant step size. - **X (Evaluation iterate)**: A running average that represents the "best" model. - **Y (Interpolation point)**: Where the gradient is calculated. The update includes **Explicit Momentum** to smooth the aggressive spectral steps and **Row-wise Normalization** to ensure per-neuron stability. ## Experiments: Closing the SOTA Gap The authors tested SF-NorMuon on LLaMA-style transformers (125M and 772M parameters) using the FineWeb-100B dataset. - **Speedup**: SF-NorMuon offers a massive efficiency gain (up to 52% faster than SF-AdamW to reach the same loss). - **Robustness**: It maintains high performance across a wide range of learning rates, whereas standard AdamW is brittle. - **Long Horizon**: While other methods crashed after the standard "Chinchilla" limit, SF-NorMuon remained stable up to 48x Chinchilla ratios.  *Figure 3: Ablation studies confirm that explicit momentum (smoothing the gradient before polar decomposition) is vital for spectral anytime training.* ## Conclusion: A Step Toward Continual Learning SF-NorMuon is more than just a better optimizer; it is a shift in mindset. By decoupling optimization from time, we can treat LLM training as an open-ended process. You can pull a SOTA-quality checkpoint at any billion-token interval without ever looking at a decay curve. **Key Values for Practitioners:** - **Zero-tuning for horizons**: Use the same hyperparameters whether you train for 1 day or 1 month. - **Spectral Efficiency**: Get the convergence benefits of Muon/NorMuon without the baggage of schedules. - **Stability**: Mathematical guarantees of boundedness through smart weight decay placement.