Data is the Destiny of Generalization: Decoupling Loss-to-Loss Scaling Laws

Llms on the line: Data determines loss-to-loss scaling laws

2025-01-01
Prasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge, Wieland Brendel
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates "loss-to-loss" scaling laws in Large Language Models (LLMs), which relate pretraining losses to downstream task performance. Through an extensive interventional analysis of over 6,000 checkpoints, the authors demonstrate that pretraining data is the primary determinant of these scaling trends, whereas architecture (e.g., Transformer vs. Mamba) and hyperparameters have surprisingly minimal impact.

TL;DR

For years, the AI community has obsessed over architectural innovations (Transformer vs. SSMs) and hyperparameter tuning. This paper delivers a sobering empirical reality check: Data determines the generalization curve. If two models—even those as fundamentally different as Llama (Transformer) and Mamba (State-Space Model)—reach the same pretraining loss on the same data, their downstream task performance will be nearly identical.

The "Accuracy on the Line" Intuition

In computer vision, researchers previously discovered that a model’s performance on out-of-distribution (OOD) data is linearly correlated with its in-distribution accuracy. This paper brings that "on the line" philosophy to the world of LLMs.

The authors argue that we should stop looking at scaling solely as "Compute Training Loss" and instead focus on the Loss-to-Loss relationship: How does an incremental improvement in pretraining perplexity translate to an improvement in solving a logic puzzle (ARC) or finishing a sentence (Hellaswag)?

Methodology: The Interventional Approach

To find the "root cause" of generalization, the authors analyzed over 6,000 model checkpoints. They used a shifted power law to model the relationship: Where is the pretraining loss and is the downstream task loss. By changing one variable at a time (e.g., switching the optimizer but keeping the data same), they measured how much the resulting curve "shifted."

Model Architecture and Factor Analysis Figure 1: The core finding: Pretraining data is the only factor that significantly moves the needle on the loss-to-loss scaling trend.

Key Findings

1. Architecture is "Generalization-Neutral"

Perhaps the most shocking result: Architecture has almost no impact. The study compared Llama (Attention-based) and Mamba (SSM-based). Despite these models having fundamentally different inductive biases—one using global attention and the other using structured state-space models—they fall on the same loss-to-loss curve when trained on the same data.

Architecture Intervention Results Figure 2: Llama and Mamba checkpoints following nearly identical scaling paths.

2. Data is the "Master Control"

When the researchers swapped the pretraining dataset (e.g., from FineWeb-Edu to C4), the scaling laws shifted significantly. This suggests that the "intrinsic difficulty" and "transferability" of knowledge are baked into the data distribution itself. Deduplicating data (The Pile vs. The Pile Deduped) even caused measurable shifts, proving that distribution matters more than model design.

Data Intervention Results Figure 3: Changing pretraining data causes a massive "jump" in the generalization curve.

3. Hyperparameters and Size: Just Points on the Same Line

Whether you use AdamW or a WSD schedule, or whether your model has 70M or 7B parameters, you are simply moving along the same line defined by your data. Larger models aren't "smarter" in how they generalize; they are just further along the curve because they can achieve lower training loss.

Critical Analysis & Takeaways

  • The "Inductive Bias" Illusion: We often claim new architectures provide better "world models." This study suggests that modern LLM architectures are essentially commodities. Their primary value lies in training efficiency (FLOPs per token), not in their ability to generalize differently from a given set of data.
  • Practical Advice: If your goal is to beat a benchmark, don't waste months designing a new layer. Invest that time in curating a higher-quality pretraining corpus. Data is the ceiling for your model's downstream potential.
  • Limitations: The study focuses on zero-shot generalization. It remains to be seen if these laws hold after extensive supervised fine-tuning (SFT) or RLHF, where the model might deviate from the pretraining distribution's "line."

Conclusion

This work reinforces a growing sentiment in the industry: Data is the moat. By showing that architecture, size, and optimization are largely subservient to the data distribution, the authors provide a rigorous empirical foundation for the "Data-Centric AI" movement in the era of LLMs.

Find Similar Papers

Try Our Examples

  • Search for recent papers that investigate how pretraining data diversity or quality specifically affects the power-law coefficients of downstream scaling.
  • Which paper first introduced the "loss-to-loss" scaling law framework, and how does this study's treatment of irreducible error (E) differ from that original work?
  • Are there any studies exploring if "loss-to-loss" scaling laws remain consistent in multi-modal (Vision-Language) models or Reinforcement Learning agents?
Contents
Data is the Destiny of Generalization: Decoupling Loss-to-Loss Scaling Laws
1. TL;DR
2. The "Accuracy on the Line" Intuition
3. Methodology: The Interventional Approach
4. Key Findings
4.1. 1. Architecture is "Generalization-Neutral"
4.2. 2. Data is the "Master Control"
4.3. 3. Hyperparameters and Size: Just Points on the Same Line
5. Critical Analysis & Takeaways
6. Conclusion