OPUS: Breaking the Data Wall via Optimizer-Aware Selection

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

2026-02-01
Shaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu, Jialin Liu, Guo Chen, Tianyu Zhang, Junhao Zheng, Kexin Yang, Xingzhang Ren, Dayiheng Liu, Linfeng Zhang
Summary
Problem
Method
Results
Takeaways
Abstract

OPUS (Optimizer-induced Projected Utility Selection) is a dynamic data selection framework for LLM pre-training that scores data candidates based on their utility within the specific update geometry of modern optimizers like AdamW and Muon. It utilizes "ghost" gradients and CountSketch projections to achieve state-of-the-art data efficiency, outperforming industrial baselines and even 6.6x longer training runs with minimal (4.7%) compute overhead.

TL;DR

As high-quality data becomes a scarce commodity, the focus of LLM pre-training is shifting from "more tokens" to "better tokens." OPUS (Optimizer-induced Projected Utility Selection) is a novel framework that makes data selection "optimizer-aware." By calculating data utility within the actual geometric space of optimizers like AdamW and Muon, OPUS achieves massive efficiency gains—often outperforming models trained on twice as much data while adding less than 5% compute overhead.

The Problem: The Geometry Mismatch in Data Selection

Most current pre-training pipelines use static filters (like FineWeb-Edu) or dynamic selectors that look at raw gradients. However, modern LLMs aren't trained with simple SGD; they use adaptive optimizers (AdamW, Muon) that precondition and reshape the gradient.

If your selection criteria assumes a straight-line path (SGD) while your optimizer is driving on a curved manifold (AdamW), you select tokens that are technically "high signal" but practically useless for the actual step the model is about to take. This mismatch leads to inefficient training and slower convergence.

Methodology: High-Fidelity Utility Estimation

OPUS solves this by defining Optimizer-induced Utility. The core intuition is: A batch is only valuable if it moves the model parameters in a direction that improves performance on a target distribution, under the optimizer's specific geometry.

1. Principled Utility Objective

OPUS calculates a score based on how well a candidate's preconditioned update aligns with a "Proxy" direction (representing high-quality knowledge). It also includes a Redundancy Penalty (an isotropic Hessian approximation) to ensure the selected batch is diverse and doesn't just repeat the same signal.

2. Scalability via Ghost & Sketch

Calculating per-sample gradients for billions of parameters is impossible. OPUS uses two clever tricks:

  • Ghost Technique: Exploits the rank-1 structure of gradients in linear layers to avoid full materialization.
  • CountSketch: Projects these high-dimensional updates into a low-dimensional space () to compute inner products efficiently.

OPUS Architecture and Workflow Figure: The OPUS pipeline integrating BENCH-PROXY retrieval, projected scoring, and Boltzmann sampling.

Experiments: More with Less

The authors tested OPUS across GPT-2 scales and the Qwen3-8B model. The results are striking.

From-Scratch Pre-training

On the FineWeb dataset, OPUS-trained models consistently outperformed industrial-grade static filters and prior dynamic selectors.

  • GPT-2 XL Performance: OPUS achieved an average accuracy of 41.75% across benchmarks, whereas random sampling required double the tokens (60B vs 30B) to reach similar levels (41.29%).
  • Efficiency: It achieved an 8x reduction in computation for GPT-XL on FineWeb targets.

Performance Comparison Figure: OPUS outperforms random selection and prior baselines across multiple benchmarks.

Continued Pre-training (CPT)

When adapting Qwen3-8B to specialized science domains (SciencePedia), OPUS demonstrated its "super-power." It reached peak performance using only 0.5B tokens, already surpassing a random selection baseline trained on 3B tokens—a 6x efficiency gain.

CPT Results Figure: Convergence speed on SciencePedia for Qwen3-8B.

Deep Insight: Why it Works

The "Secret Sauce" of OPUS is its Adaptive Preconditioner.

  • In AdamW, it rescales coordinates based on the inverse square root of the second moment.
  • In Muon, it accounts for Newton-Schulz orthogonalization.

By "freezing" these preconditioners within a single step, OPUS can rank thousands of documents per second to find the ones that best match the model's current "learning appetite." The use of Boltzmann Sampling instead of greedy top-k selection prevents the model from collapsing into a narrow distribution of "easy" educational tokens, maintaining the general-purpose capabilities of the LLM.

Conclusion

OPUS demonstrates that the path to better LLMs isn't just about scraping more of the web—it's about the math of selection. By aligning data utility with optimizer dynamics, we can extract significantly more intelligence per FLOP.

Limitations: The framework currently relies on a "Proxy" set (BENCH-PROXY). If the proxy quality is poor or biased, the selection will follow. Future work remains to be done in making the proxy generation entirely self-supervised or automatically balanced across all human knowledge.

Find Similar Papers

Try Our Examples

  • Examine recent papers that optimize data selection by tracking the training trajectory or influence of samples over time in LLMs.
  • What are the theoretical foundations of the "Ghost technique" for per-sample gradient estimation, and how has it been applied in previous dynamic selection frameworks like GREATS?
  • Explore research applying optimizer-aware or geometry-aligned data selection to multimodal models or reinforcement learning from human feedback (RLHF).
Contents
OPUS: Breaking the Data Wall via Optimizer-Aware Selection
1. TL;DR
2. The Problem: The Geometry Mismatch in Data Selection
3. Methodology: High-Fidelity Utility Estimation
3.1. 1. Principled Utility Objective
3.2. 2. Scalability via Ghost & Sketch
4. Experiments: More with Less
4.1. From-Scratch Pre-training
4.2. Continued Pre-training (CPT)
5. Deep Insight: Why it Works
6. Conclusion