One Learning Rate Doesn't Fit All: Balancing LLM Training with Heavy-Tail Guidance

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs

2026-05-01
Di He, Songjun Tu, Keyu Wang, Lu Yin, Shiwei Liu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Layerwise Learning Rate (LLR), an adaptive training scheme for Large Language Models (LLMs) that assigns distinct learning rates to individual Transformer layers based on Heavy-Tailed Self-Regularization (HT-SR) theory. LLR achieves up to a 1.5× training speedup and significant improvements in zero-shot accuracy (+2% for 1B/3B models) across various architectures and optimizers.

Executive Summary

TL;DR: The era of applying a single, uniform learning rate (LR) to every layer of a Transformer might be coming to an end. Researchers have introduced Layerwise Learning Rate (LLR), a strategy that dynamically adjusts the LR for different layers by measuring the "Heavy-Tailedness" of their weight matrices. By accelerating the training of lagging components (FFNs) and refining well-trained ones (Attention), LLR achieves a 1.5× speedup and significant gains in model intelligence (up to +2.0% zero-shot accuracy) with virtually zero extra tuning cost.

In the landscape of LLM optimization, this work shifts from "generic scaling" to "structural sensitivity," proving that understanding the internal spectral health of a model is key to unlocking its full potential.

The "Uniformity" Trap

For years, we have treated Transformers as homogeneous blocks. Whether it's the Embedding layer, the Multi-Head Attention, or the Feed-Forward Network (FFN), they all usually share the same global learning rate.

However, these components have vastly different mathematical properties and "learning speeds." Previous attempts to fix this, such as LARS or LAMB, were built for the CNN era. When applied to LLMs, they often struggle to beat a simple, well-tuned AdamW baseline. The core problem is: how do we know which layer needs more "push" and which needs "restraint" without doing a thousand grid searches?

Methodology: The "Stethoscope" for Weights

The authors find the answer in Heavy-Tailed Self-Regularization (HT-SR) theory.

1. Measuring Spectral Health

Think of the Empirical Spectral Density (ESD) of a weight matrix as its "training pulse." In HT-SR theory, a "heavy tail" in the spectrum (a power-law distribution) indicates that the layer has learned strong, meaningful correlations.

The degree of this heavy-tailedness is measured by the PL_Alpha_Hill () metric:

  • Low : Strong heavy-tail. The layer is "well-trained."
  • High : Weak tail. The layer is "under-trained."

2. The LLR Mapping

LLR calculates these values for every layer every few hundred steps. It then maps them to a specific LR:

  • Layers with high (FFNs, Embeddings) get a larger LR to catch up.
  • Layers with low (Attention modules) get a smaller LR to avoid Overfitting or instability.

Model Architecture and Intuition Figure: LLR assigns higher rates to layers with higher Alpha (lagging layers) to balance the total model training.

3. Engineering for Stability

To make this work for LLMs, the authors introduced several critical "tricks":

  • Soft Layer-wise LR Switch: Avoids sudden LR spikes by linearly transitioning between the old and new rates.
  • Tailored Embedding Treatment: Keeps the embedding layer at the upper bound because it consistently shows high alpha.
  • Efficient Active Phase: Updates are only performed during the first 20% of training, where the most structural change happens, saving compute.

Experiments: Superior Gains, Faster Convergence

The results across LLaMa architectures (60M to 3B) are striking. LLR consistently outperforms Uniform LR, LARS, and LAMB.

Performance Comparison Figure: LLR is significantly less sensitive to hyperparameter tuning and outperforms other methods at their optimal settings.

Key Breakthroughs:

  • Accuracy Boost: On a 1B LLaMa model, average zero-shot accuracy rose from 47.09% to 49.02%.
  • Convergence Speed: To reach the same loss level, LLR is 1.5× faster than the standard AdamW baseline.
  • Optimizer Agnostic: It works with both AdamW and the newer Muon optimizer, suggesting the benefit is fundamental to the weight spectra, not the specific update rule.

Training Curves Figure: LLR showing significantly faster loss reduction compared to uniform baselines.

Deep Insights & Conclusion

Why does this work? By analyzing the dynamics of , the authors found that Attention parameters consistently learn faster (lower ) than FFNs. By giving FFNs a higher learning rate, we reduce the "lag" between model components.

Takeaways for the Industry:

  1. Stop using one LR: The heterogeneity of Transformers is too high to ignore.
  2. Spectral Analysis works: You don't need a validation set to judge if a layer is well-trained; its weight spectrum tells the story.
  3. Transferability: One of the best parts of LLR is that it inherits the optimal LR from a uniform baseline—you don't need to re-tune everything from scratch.

While LLR currently focuses on the first 20% of training steps, future work could explore if dynamic adjustments in the "long tail" of pre-training (trillions of tokens) could prevent the performance plateaus often seen in massive models.

Find Similar Papers

Try Our Examples

  • Look for recent papers that explore architectural heterogeneity in Transformers and its impact on optimization beyond stationary learning rates.
  • What are the foundational papers of Heavy-Tailed Self-Regularization (HT-SR) theory in deep learning, and how has its application evolved from CNNs to LLMs?
  • Identify research that applies layer-wise or block-wise scaling laws to large-scale multimodal models or vision transformers.
Contents
One Learning Rate Doesn't Fit All: Balancing LLM Training with Heavy-Tail Guidance
1. Executive Summary
2. The "Uniformity" Trap
3. Methodology: The "Stethoscope" for Weights
3.1. 1. Measuring Spectral Health
3.2. 2. The LLR Mapping
3.3. 3. Engineering for Stability
4. Experiments: Superior Gains, Faster Convergence
4.1. Key Breakthroughs:
5. Deep Insights & Conclusion
5.1. Takeaways for the Industry: