AMUSE: Stabilizing Muon for Schedule-Free Foundation Model Training
AMUSE: Anytime Muon with Stable Gradient Evaluation
AMUSE (Anytime MUon with Stable gradient Evaluation) is a novel optimization algorithm that integrates Muon's orthogonalization with Schedule-Free (SF) averaging. It achieves SOTA Pareto frontiers in LLM pretraining and vision tasks by removing the need for traditional learning rate decay schedules.
TL;DR
Optimization is a trade-off between speed and stability. Muon is fast because it "orthogonalizes" updates to ignore bad curvature, but this aggressive nature makes training unstable. AMUSE solves this by evaluating gradients at a smarter, "smoothed" location in the loss landscape. It removes the need for learning rate schedules and delivers SOTA performance in LLM pretraining and Computer Vision.
Problem & Motivation: The River and the Valley
Deep learning loss landscapes are rarely uniform. They are often described as "River Valleys":
- The Valley Walls: High-curvature directions where the loss changes rapidly. Moving here causes instability and oscillations.
- The River: A low-curvature "bulk" subspace where actual learning progress happens.
Standard Muon is great at moving along the river because its orthogonalization prevents any single direction from dominating the update. However, this same mechanism takes tiny "noise" from the valley walls and amplifies it into massive oscillations. To stop this, researchers usually use a Learning Rate Decay to force the model to settle at the end of training. But what if we want the model to be ready at any time?
Methodology: High-Stability Evaluation
The core innovation of AMUSE (Anytime MUon with Stable gradient Evaluation) is the use of a time-varying interpolation coefficient, .
The Two-Track System
AMUSE maintains two parallel sequences:
- The Fast Sequence (): Follows the raw Muon updates.
- The Averaged Sequence (): A stable, running average of previous steps that naturally stays in the "river."
Instead of calculating the gradient only at the fast point, AMUSE calculates it at an interpolated point .
The Schedule
Early in training, AMUSE looks at the fast sequence to adapt quickly ( is small). As training progresses, it shifts ((\beta_t o 1)) to evaluate gradients closer to the stable averaged sequence. This "filters" out the high-curvature noise before the orthogonalization step ever happens.
Figure 1: SGD oscillations vs. Muon's faster but still unstable progress. AMUSE aims to travel straight down the river floor.
Experiments & Results: Breaking the Pareto Frontier
AMUSE was tested against AdamW and vanilla Muon on Llama models ranging from 124M to 1B parameters, as well as ImageNet benchmarks.
1. Large Language Models (LLM)
On a 720M Llama model trained on FineWeb, AMUSE reached Muon's best performance 1.51x faster. Notably, it doesn't need a cosine decay—it stays superior throughout the entire training run.
Figure 2: Validation perplexity on FineWeb. AMUSE (purple) shows a clear lead over AdamW and Muon across different model scales.
2. Vision Tasks
In vision benchmarks like ResNet-50 and ViT fine-tuning, AMUSE consistently held the highest accuracy. In ViT fine-tuning, the speedup was a staggering 3.08x.
Critical Analysis & Conclusion
Takeaway
AMUSE proves that we don't need complex learning rate schedules to train foundation models efficiently. By mathematically aligning the optimizer's evaluation point with the "River" of the loss landscape, we get the speed of Muon with the stability of AdamW.
Limitations
The primary drawback of AMUSE is Memory Overhead. Because it tracks an extra "averaged" state (), it requires more VRAM than vanilla Muon, though it remains comparable to standard AdamW or Schedule-Free AdamW.
Future Work
The next frontier is designing "memory-efficient" versions of AMUSE that can suppress valley-wall oscillations without storing multiple copies of the model weights, making it viable for ultra-large-scale 100B+ parameter training.
