AMUSE: Stabilizing Muon for Schedule-Free Foundation Model Training

AMUSE: Anytime Muon with Stable Gradient Evaluation

2026-05-01
Jueun Kim, Baekrok Shin, Jihun Yun, Beomhan Baek, Minhak Song, Chulhee Yun
Summary
Problem
Method
Results
Takeaways
Abstract

AMUSE (Anytime MUon with Stable gradient Evaluation) is a novel optimization algorithm that integrates Muon's orthogonalization with Schedule-Free (SF) averaging. It achieves SOTA Pareto frontiers in LLM pretraining and vision tasks by removing the need for traditional learning rate decay schedules.

TL;DR

Optimization is a trade-off between speed and stability. Muon is fast because it "orthogonalizes" updates to ignore bad curvature, but this aggressive nature makes training unstable. AMUSE solves this by evaluating gradients at a smarter, "smoothed" location in the loss landscape. It removes the need for learning rate schedules and delivers SOTA performance in LLM pretraining and Computer Vision.

Problem & Motivation: The River and the Valley

Deep learning loss landscapes are rarely uniform. They are often described as "River Valleys":

  • The Valley Walls: High-curvature directions where the loss changes rapidly. Moving here causes instability and oscillations.
  • The River: A low-curvature "bulk" subspace where actual learning progress happens.

Standard Muon is great at moving along the river because its orthogonalization prevents any single direction from dominating the update. However, this same mechanism takes tiny "noise" from the valley walls and amplifies it into massive oscillations. To stop this, researchers usually use a Learning Rate Decay to force the model to settle at the end of training. But what if we want the model to be ready at any time?

Methodology: High-Stability Evaluation

The core innovation of AMUSE (Anytime MUon with Stable gradient Evaluation) is the use of a time-varying interpolation coefficient, .

The Two-Track System

AMUSE maintains two parallel sequences:

  1. The Fast Sequence (): Follows the raw Muon updates.
  2. The Averaged Sequence (): A stable, running average of previous steps that naturally stays in the "river."

Instead of calculating the gradient only at the fast point, AMUSE calculates it at an interpolated point .

The Schedule

Early in training, AMUSE looks at the fast sequence to adapt quickly ( is small). As training progresses, it shifts ((\beta_t o 1)) to evaluate gradients closer to the stable averaged sequence. This "filters" out the high-curvature noise before the orthogonalization step ever happens.

Model Architecture and River-Valley Illustration Figure 1: SGD oscillations vs. Muon's faster but still unstable progress. AMUSE aims to travel straight down the river floor.

Experiments & Results: Breaking the Pareto Frontier

AMUSE was tested against AdamW and vanilla Muon on Llama models ranging from 124M to 1B parameters, as well as ImageNet benchmarks.

1. Large Language Models (LLM)

On a 720M Llama model trained on FineWeb, AMUSE reached Muon's best performance 1.51x faster. Notably, it doesn't need a cosine decay—it stays superior throughout the entire training run.

LLM Performance Comparison Figure 2: Validation perplexity on FineWeb. AMUSE (purple) shows a clear lead over AdamW and Muon across different model scales.

2. Vision Tasks

In vision benchmarks like ResNet-50 and ViT fine-tuning, AMUSE consistently held the highest accuracy. In ViT fine-tuning, the speedup was a staggering 3.08x.

Critical Analysis & Conclusion

Takeaway

AMUSE proves that we don't need complex learning rate schedules to train foundation models efficiently. By mathematically aligning the optimizer's evaluation point with the "River" of the loss landscape, we get the speed of Muon with the stability of AdamW.

Limitations

The primary drawback of AMUSE is Memory Overhead. Because it tracks an extra "averaged" state (), it requires more VRAM than vanilla Muon, though it remains comparable to standard AdamW or Schedule-Free AdamW.

Future Work

The next frontier is designing "memory-efficient" versions of AMUSE that can suppress valley-wall oscillations without storing multiple copies of the model weights, making it viable for ultra-large-scale 100B+ parameter training.

Find Similar Papers

Try Our Examples

  • Find recent papers other than AMUSE that attempt to solve the valley-wall oscillation problem in Muon or other orthogonalization-based optimizers.
  • Which paper first introduced the "river-valley" loss landscape perspective, and how does AMUSE's implementation of this theory differ from previous "subspace-aware" optimizers?
  • Explore if there are studies applying the AMUSE interpolation mechanism to reinforcement learning or audio generation tasks where training stability is traditionally a bottleneck.
Contents
AMUSE: Stabilizing Muon for Schedule-Free Foundation Model Training
1. TL;DR
2. Problem & Motivation: The River and the Valley
3. Methodology: High-Stability Evaluation
3.1. The Two-Track System
3.2. The $\beta_t$ Schedule
4. Experiments & Results: Breaking the Pareto Frontier
4.1. 1. Large Language Models (LLM)
4.2. 2. Vision Tasks
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Work