A Bitter Lesson for Data Filtering: Is the Best Filter No Filter?
A Bitter Lesson for Data Filtering
This paper presents a large-scale scaling study on data filtering for LLMs, introducing the "Bitter Lesson" for data curation: as compute and model size scale, the optimal strategy converges toward using no filtering at all. By training models up to 7B parameters on raw vs. filtered Common Crawl, the authors demonstrate that sufficiently large models eventually outperform filtered baselines by extracting signal from nominally "poor" data.
TL;DR
Rich Sutton's "Bitter Lesson" famously argued that general methods that leverage compute eventually win over human-authored heuristics. This paper applies that logic to data curation. The core finding: Sufficiently large LLMs are remarkably "noise-tolerant." Given enough compute, training on the raw, "messy" Common Crawl outperforms even the most sophisticated filtered datasets like RefinedWeb or DCLM.
Context: The Data Scarcity vs. Quality Trade-off
For years, the industry standard has been: Clean the data until only the "gold" remains. This led to datasets like DCLM-Baseline, which discards 99% of the web to focus on quality. But as models grow toward 1T parameters, we are running out of "gold" data.
The authors ask a radical question: What if the "junk" we are throwing away actually contains useful signals that only large models can see?
The Core Insight: The Crossing Point ()
The researchers discovered that "quality" is not an intrinsic property of data, but a relationship between data and model capacity.
- Small Models (15M - 80M): These models are easily "confused" by noise. Filtering is essential for them because they lack the capacity to distinguish between high-signal text and gibberish.
- Large Models (330M - 7B): These models act as robust statistical filters themselves. They can extract structural information even from word-shuffled text.
Methodology: Scaling the Pool
The authors tested five major filter types (English, Repetition, Stop Words, RefinedWeb, and DCLM-Baseline) against a raw Common Crawl pool.
Figure 1: Notice how the black line (Unfiltered Pool) starts as the worst performer but crosses almost every filtered version as training steps increase.
Experiments in Content Destruction: Junk Data Injection
To push the limits, the authors didn't just use "bad" web data; they created hallucinatory junk:
- Random Strings: Entirely synthetic gibberish.
- Shuffled Documents: Real CC documents where the word order was randomized (destroying syntax but keeping unigram/vocabulary frequency).
The shocker: The 330M+ models actually benefited from the shuffled data. This suggests that even without syntax, the mere co-occurrence of words (e.g., "Paris" appearing near "France" in a shuffled pile) provides enough "unigram juice" for a large model to improve its world knowledge.
Figure 2: Large models show near-complete recovery from random noise and actual gains from shuffled text.
Scaling Laws for the Future
When will "No Filter" become the industry standard? The authors projected a scaling law for the 240-Trillion-token DCLM-Pool.
They predict the "Crossing Point" (where raw CC beats RefinedWeb) will occur at FLOPs. While current frontier models (like Grok-3 or GPT-4) are estimated at , exponential growth in compute suggests we will hit the "No Filter" era by 2030.

Why Does This Work? A Linear Intuition
The authors provide a mathematical proof using Low-Rank Matrix Factorization. Think of the model as a matrix trying to learn different "tasks" (languages, topics).
- If the model's Rank (Capacity) is low, different tasks interfere with each other—noise drowns out signal.
- If the model's Rank is high enough, it can assign separate "subspaces" to the signal and the noise. The noise simply occupies an orthogonal dimension, leaving the signal intact.
Critical Analysis & Conclusion
Limitations
- Factuality: While models are robust to noise (shuffled words), they may still be vulnerable to misinformation (false facts), which acts as a "correctly labeled but wrong" signal.
- Architecture: These results are for Dense Transformers. Whether sparse models like MoE (Mixture of Experts) handle noise as gracefully is still an open question.
Final Takeaway
The era of artisanal data cleaning may be coming to an end. This paper suggests that instead of spending millions on sophisticated BERT-based classifiers to prune our data, we should spend those millions on GPUs to scale our models. The "messiness" of the internet isn't a bug—it's a feature that only the largest models are smart enough to use.
