Data is the Destiny of Generalization: Decoupling Loss-to-Loss Scaling Laws
Llms on the line: Data determines loss-to-loss scaling laws
This paper investigates "loss-to-loss" scaling laws in Large Language Models (LLMs), which relate pretraining losses to downstream task performance. Through an extensive interventional analysis of over 6,000 checkpoints, the authors demonstrate that pretraining data is the primary determinant of these scaling trends, whereas architecture (e.g., Transformer vs. Mamba) and hyperparameters have surprisingly minimal impact.
TL;DR
For years, the AI community has obsessed over architectural innovations (Transformer vs. SSMs) and hyperparameter tuning. This paper delivers a sobering empirical reality check: Data determines the generalization curve. If two models—even those as fundamentally different as Llama (Transformer) and Mamba (State-Space Model)—reach the same pretraining loss on the same data, their downstream task performance will be nearly identical.
The "Accuracy on the Line" Intuition
In computer vision, researchers previously discovered that a model’s performance on out-of-distribution (OOD) data is linearly correlated with its in-distribution accuracy. This paper brings that "on the line" philosophy to the world of LLMs.
The authors argue that we should stop looking at scaling solely as "Compute Training Loss" and instead focus on the Loss-to-Loss relationship: How does an incremental improvement in pretraining perplexity translate to an improvement in solving a logic puzzle (ARC) or finishing a sentence (Hellaswag)?
Methodology: The Interventional Approach
To find the "root cause" of generalization, the authors analyzed over 6,000 model checkpoints. They used a shifted power law to model the relationship: Where is the pretraining loss and is the downstream task loss. By changing one variable at a time (e.g., switching the optimizer but keeping the data same), they measured how much the resulting curve "shifted."
Figure 1: The core finding: Pretraining data is the only factor that significantly moves the needle on the loss-to-loss scaling trend.
Key Findings
1. Architecture is "Generalization-Neutral"
Perhaps the most shocking result: Architecture has almost no impact. The study compared Llama (Attention-based) and Mamba (SSM-based). Despite these models having fundamentally different inductive biases—one using global attention and the other using structured state-space models—they fall on the same loss-to-loss curve when trained on the same data.
Figure 2: Llama and Mamba checkpoints following nearly identical scaling paths.
2. Data is the "Master Control"
When the researchers swapped the pretraining dataset (e.g., from FineWeb-Edu to C4), the scaling laws shifted significantly. This suggests that the "intrinsic difficulty" and "transferability" of knowledge are baked into the data distribution itself. Deduplicating data (The Pile vs. The Pile Deduped) even caused measurable shifts, proving that distribution matters more than model design.
Figure 3: Changing pretraining data causes a massive "jump" in the generalization curve.
3. Hyperparameters and Size: Just Points on the Same Line
Whether you use AdamW or a WSD schedule, or whether your model has 70M or 7B parameters, you are simply moving along the same line defined by your data. Larger models aren't "smarter" in how they generalize; they are just further along the curve because they can achieve lower training loss.
Critical Analysis & Takeaways
- The "Inductive Bias" Illusion: We often claim new architectures provide better "world models." This study suggests that modern LLM architectures are essentially commodities. Their primary value lies in training efficiency (FLOPs per token), not in their ability to generalize differently from a given set of data.
- Practical Advice: If your goal is to beat a benchmark, don't waste months designing a new layer. Invest that time in curating a higher-quality pretraining corpus. Data is the ceiling for your model's downstream potential.
- Limitations: The study focuses on zero-shot generalization. It remains to be seen if these laws hold after extensive supervised fine-tuning (SFT) or RLHF, where the model might deviate from the pretraining distribution's "line."
Conclusion
This work reinforces a growing sentiment in the industry: Data is the moat. By showing that architecture, size, and optimization are largely subservient to the data distribution, the authors provide a rigorous empirical foundation for the "Data-Centric AI" movement in the era of LLMs.
