OPUS: Breaking the Data Wall via Optimizer-Aware Selection
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
OPUS (Optimizer-induced Projected Utility Selection) is a dynamic data selection framework for LLM pre-training that scores data candidates based on their utility within the specific update geometry of modern optimizers like AdamW and Muon. It utilizes "ghost" gradients and CountSketch projections to achieve state-of-the-art data efficiency, outperforming industrial baselines and even 6.6x longer training runs with minimal (4.7%) compute overhead.
TL;DR
As high-quality data becomes a scarce commodity, the focus of LLM pre-training is shifting from "more tokens" to "better tokens." OPUS (Optimizer-induced Projected Utility Selection) is a novel framework that makes data selection "optimizer-aware." By calculating data utility within the actual geometric space of optimizers like AdamW and Muon, OPUS achieves massive efficiency gains—often outperforming models trained on twice as much data while adding less than 5% compute overhead.
The Problem: The Geometry Mismatch in Data Selection
Most current pre-training pipelines use static filters (like FineWeb-Edu) or dynamic selectors that look at raw gradients. However, modern LLMs aren't trained with simple SGD; they use adaptive optimizers (AdamW, Muon) that precondition and reshape the gradient.
If your selection criteria assumes a straight-line path (SGD) while your optimizer is driving on a curved manifold (AdamW), you select tokens that are technically "high signal" but practically useless for the actual step the model is about to take. This mismatch leads to inefficient training and slower convergence.
Methodology: High-Fidelity Utility Estimation
OPUS solves this by defining Optimizer-induced Utility. The core intuition is: A batch is only valuable if it moves the model parameters in a direction that improves performance on a target distribution, under the optimizer's specific geometry.
1. Principled Utility Objective
OPUS calculates a score based on how well a candidate's preconditioned update aligns with a "Proxy" direction (representing high-quality knowledge). It also includes a Redundancy Penalty (an isotropic Hessian approximation) to ensure the selected batch is diverse and doesn't just repeat the same signal.
2. Scalability via Ghost & Sketch
Calculating per-sample gradients for billions of parameters is impossible. OPUS uses two clever tricks:
- Ghost Technique: Exploits the rank-1 structure of gradients in linear layers to avoid full materialization.
- CountSketch: Projects these high-dimensional updates into a low-dimensional space () to compute inner products efficiently.
Figure: The OPUS pipeline integrating BENCH-PROXY retrieval, projected scoring, and Boltzmann sampling.
Experiments: More with Less
The authors tested OPUS across GPT-2 scales and the Qwen3-8B model. The results are striking.
From-Scratch Pre-training
On the FineWeb dataset, OPUS-trained models consistently outperformed industrial-grade static filters and prior dynamic selectors.
- GPT-2 XL Performance: OPUS achieved an average accuracy of 41.75% across benchmarks, whereas random sampling required double the tokens (60B vs 30B) to reach similar levels (41.29%).
- Efficiency: It achieved an 8x reduction in computation for GPT-XL on FineWeb targets.
Figure: OPUS outperforms random selection and prior baselines across multiple benchmarks.
Continued Pre-training (CPT)
When adapting Qwen3-8B to specialized science domains (SciencePedia), OPUS demonstrated its "super-power." It reached peak performance using only 0.5B tokens, already surpassing a random selection baseline trained on 3B tokens—a 6x efficiency gain.
Figure: Convergence speed on SciencePedia for Qwen3-8B.
Deep Insight: Why it Works
The "Secret Sauce" of OPUS is its Adaptive Preconditioner.
- In AdamW, it rescales coordinates based on the inverse square root of the second moment.
- In Muon, it accounts for Newton-Schulz orthogonalization.
By "freezing" these preconditioners within a single step, OPUS can rank thousands of documents per second to find the ones that best match the model's current "learning appetite." The use of Boltzmann Sampling instead of greedy top-k selection prevents the model from collapsing into a narrow distribution of "easy" educational tokens, maintaining the general-purpose capabilities of the LLM.
Conclusion
OPUS demonstrates that the path to better LLMs isn't just about scraping more of the web—it's about the math of selection. By aligning data utility with optimizer dynamics, we can extract significantly more intelligence per FLOP.
Limitations: The framework currently relies on a "Proxy" set (BENCH-PROXY). If the proxy quality is poor or biased, the selection will follow. Future work remains to be done in making the proxy generation entirely self-supervised or automatically balanced across all human knowledge.
