One Learning Rate Doesn't Fit All: Balancing LLM Training with Heavy-Tail Guidance
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
The paper introduces Layerwise Learning Rate (LLR), an adaptive training scheme for Large Language Models (LLMs) that assigns distinct learning rates to individual Transformer layers based on Heavy-Tailed Self-Regularization (HT-SR) theory. LLR achieves up to a 1.5× training speedup and significant improvements in zero-shot accuracy (+2% for 1B/3B models) across various architectures and optimizers.
Executive Summary
TL;DR: The era of applying a single, uniform learning rate (LR) to every layer of a Transformer might be coming to an end. Researchers have introduced Layerwise Learning Rate (LLR), a strategy that dynamically adjusts the LR for different layers by measuring the "Heavy-Tailedness" of their weight matrices. By accelerating the training of lagging components (FFNs) and refining well-trained ones (Attention), LLR achieves a 1.5× speedup and significant gains in model intelligence (up to +2.0% zero-shot accuracy) with virtually zero extra tuning cost.
In the landscape of LLM optimization, this work shifts from "generic scaling" to "structural sensitivity," proving that understanding the internal spectral health of a model is key to unlocking its full potential.
The "Uniformity" Trap
For years, we have treated Transformers as homogeneous blocks. Whether it's the Embedding layer, the Multi-Head Attention, or the Feed-Forward Network (FFN), they all usually share the same global learning rate.
However, these components have vastly different mathematical properties and "learning speeds." Previous attempts to fix this, such as LARS or LAMB, were built for the CNN era. When applied to LLMs, they often struggle to beat a simple, well-tuned AdamW baseline. The core problem is: how do we know which layer needs more "push" and which needs "restraint" without doing a thousand grid searches?
Methodology: The "Stethoscope" for Weights
The authors find the answer in Heavy-Tailed Self-Regularization (HT-SR) theory.
1. Measuring Spectral Health
Think of the Empirical Spectral Density (ESD) of a weight matrix as its "training pulse." In HT-SR theory, a "heavy tail" in the spectrum (a power-law distribution) indicates that the layer has learned strong, meaningful correlations.
The degree of this heavy-tailedness is measured by the PL_Alpha_Hill () metric:
- Low : Strong heavy-tail. The layer is "well-trained."
- High : Weak tail. The layer is "under-trained."
2. The LLR Mapping
LLR calculates these values for every layer every few hundred steps. It then maps them to a specific LR:
- Layers with high (FFNs, Embeddings) get a larger LR to catch up.
- Layers with low (Attention modules) get a smaller LR to avoid Overfitting or instability.
Figure: LLR assigns higher rates to layers with higher Alpha (lagging layers) to balance the total model training.
3. Engineering for Stability
To make this work for LLMs, the authors introduced several critical "tricks":
- Soft Layer-wise LR Switch: Avoids sudden LR spikes by linearly transitioning between the old and new rates.
- Tailored Embedding Treatment: Keeps the embedding layer at the upper bound because it consistently shows high alpha.
- Efficient Active Phase: Updates are only performed during the first 20% of training, where the most structural change happens, saving compute.
Experiments: Superior Gains, Faster Convergence
The results across LLaMa architectures (60M to 3B) are striking. LLR consistently outperforms Uniform LR, LARS, and LAMB.
Figure: LLR is significantly less sensitive to hyperparameter tuning and outperforms other methods at their optimal settings.
Key Breakthroughs:
- Accuracy Boost: On a 1B LLaMa model, average zero-shot accuracy rose from 47.09% to 49.02%.
- Convergence Speed: To reach the same loss level, LLR is 1.5× faster than the standard AdamW baseline.
- Optimizer Agnostic: It works with both AdamW and the newer Muon optimizer, suggesting the benefit is fundamental to the weight spectra, not the specific update rule.
Figure: LLR showing significantly faster loss reduction compared to uniform baselines.
Deep Insights & Conclusion
Why does this work? By analyzing the dynamics of , the authors found that Attention parameters consistently learn faster (lower ) than FFNs. By giving FFNs a higher learning rate, we reduce the "lag" between model components.
Takeaways for the Industry:
- Stop using one LR: The heterogeneity of Transformers is too high to ignore.
- Spectral Analysis works: You don't need a validation set to judge if a layer is well-trained; its weight spectrum tells the story.
- Transferability: One of the best parts of LLR is that it inherits the optimal LR from a uniform baseline—you don't need to re-tune everything from scratch.
While LLR currently focuses on the first 20% of training steps, future work could explore if dynamic adjustments in the "long tail" of pre-training (trillions of tokens) could prevent the performance plateaus often seen in massive models.
