The Shannon Scaling Law: LLMs as Noisy Channels and the End of Monotonic Scaling

LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws

2026-05-01
Xu Ouyang, Deyi Liu, Yuhang Cai, Jing Liu, Yuan Yang, Chen Zheng, Thomas Hartvigsen, Yiyuan Ma
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the Shannon Scaling Law, a unified framework that models LLM training as information transmission over a noisy channel based on the Shannon–Hartley theorem. It maps model size to bandwidth and tokens to signal power, consistently outperforming traditional power laws in predicting performance across diverse regimes including quantization, overtraining, and noise injection.

TL;DR

Scaling laws have long been the North Star of AI, but the "bigger is always better" mantra is hitting a wall. A new paper, LLMs as Noisy Channels, introduces the Shannon Scaling Law. By treating LLMs as communication channels, the authors replace simple power laws with information theory. This new law successfully models why performance sometimes drops as you scale (catastrophic overtraining/quantization) and accurately predicts model behavior even 1.7x beyond the training data limit.

Background Positioning

In the academic coordinate system, this work moves LLM theory from Empirical Observation (fitting curves to data) to Theoretical Grounding (applying Shannon’s Information Theory). While OpenAI’s original law and DeepSeek’s Chinchilla law work in "clean" regimes, they collapse under noise. This paper provides a unified theory for both standard pretraining and "messy" real-world scenarios like low-bit quantization and noisy fine-tuning.


1. The Death of Monotonicity: Why Scaling Breaks

The AI industry has been obsessed with the idea that . However, we are seeing more "U-shaped" performance curves:

  1. Catastrophic Overtraining: Training too long on a specific task (SFT) can actually make the model worse.
  2. Quantization Sensitivity: Larger models are sometimes more fragile when compressed to 2-bit or 3-bit.

Traditional laws can't see these "basins" of loss because they don't understand Noise. The authors argue that as you scale (Parameters) and (Tokens), you aren't just scaling knowledge—you are scaling the inherent noise of the system.


2. Methodology: From Hertz to Hyperparameters

The core innovation is mapping the Shannon-Hartley Theorem () to LLM training dynamics:

  • Bandwidth (): Mapped to Model Size (). A larger model has more "frequencies" to capture complex patterns.
  • Signal (): Mapped to Training Tokens (). More data provides the signal to be transmitted into the weights.
  • Noise (): Decomposed into Data noise (typos/errors), Model noise (random initialization/interference), and Irreducible noise (architectural limits).

The Math of Performance

The loss is defined as the reciprocal of this capacity:

Model Architecture and Analogy Figure: The structural correspondence between the Shannon communication model and LLMs.


3. The "Loss Basin" and Empirical Evidence

The researchers tested this on Pythia and OLMo2 across three "perturbation" types: Gaussian noise, SFT, and Quantization.

The Visualization of Scaling Failure

In the high-SNR (clean) regime, the loss landscape is open. But as noise increases, a "basin" forms. If you move too far left (too small) or too far right (overtrained), the loss shoots up.

Loss Contours under Noise Figure: Evolution of Pythia loss contours. Note the transition from monotonic improvement to U-shaped degradation at 10dB noise.

Results: Predictive Power

The most impressive feat is Extrapolation. Standard laws fail when asked to predict the performance of a model larger than any seen in the training set. The Shannon Scaling Law maintained an of 0.847 when predicting a 12B model from 6.9B data, while OpenAI and Chinchilla laws essentially collapsed.


4. Why This Works: The Exponent Inversion

The authors conducted an ablation study on the exponents (). They found a critical insight:

  • In Low-SNR regimes, the Noise exponent () grows faster than the Bandwidth exponent ().
  • Conclusion: Scaling model size becomes detrimental because the noise amplification outpaces the information gain.

This explains the "Over-parameterization Paradox"—if your data signal isn't clean enough, a bigger model will just "memorize" the noise more efficiently, leading to worse generalization.


5. Critical Analysis & Future Outlook

Takeaways

  • Stop Brute-Force Scaling: We are approaching a regime where data quality (SNR) is more important than raw token count.
  • Strategic Training: This law helps practitioners find the "Sweet Spot"—the bottom of the loss basin—before degradation begins.

Limitations

  • Constant Fitting: The law requires fitting 6-9 parameters, which may require several small-scale "pilot" runs (N x D grid) to identify the specific noise profile of a dataset.
  • Architecture Specificity: While tested on Transformers (Pythia/OLMo2), it remains to be seen if modern architectures like Mamba or RWKV follow the exact same noise-interaction () scaling.

The Shannon Scaling Law provides a much-needed theoretical rug to stand on as we push models into the trillion-parameter range where empirical observation alone is no longer enough to avoid catastrophic failure.

Find Similar Papers

Try Our Examples

  • Search for recent papers investigating non-monotonic scaling behaviors in LLMs, specifically focusing on "catastrophic overtraining" and its relationship to downstream performance.
  • Analyze the origins of "Model-Interaction Noise" in deep learning theory—did earlier Information Bottleneck principle papers predict performance degradation when model capacity exceeds data signal?
  • Examine how the Shannon Scaling Law could be applied to Multi-modal LLMs or Reinforcement Learning, where the signal-to-noise ratio of training data (e.g., video or sparse rewards) is significantly lower than text.
Contents
The Shannon Scaling Law: LLMs as Noisy Channels and the End of Monotonic Scaling
1. TL;DR
2. Background Positioning
3. 1. The Death of Monotonicity: Why Scaling Breaks
4. 2. Methodology: From Hertz to Hyperparameters
4.1. The Math of Performance
5. 3. The "Loss Basin" and Empirical Evidence
5.1. The Visualization of Scaling Failure
5.2. Results: Predictive Power
6. 4. Why This Works: The Exponent Inversion
7. 5. Critical Analysis & Future Outlook
7.1. Takeaways
7.2. Limitations