RL Excursions: Shaking the Foundations of the LLM Training Pipeline

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

2026-06-01
Rachit Bansal, Clara Mohri, Tian Qin, David Alvarez-Melis, Sham Kakade
Summary
Problem
Method
Results
Takeaways
Abstract

The paper investigates the application of Reinforcement Learning (RL) during the early stages of LLM pre-training, rather than as a final post-training step. Using a 1B parameter model, the authors demonstrate that "Direct RL" on intermediate checkpoints (as early as 4B tokens) can match the performance of the traditional SFT→RL pipeline on reasoning tasks like GSM8K and MATH.

TL;DR

The industry-standard recipe for LLMs—Pre-train → SFT → RL—might be fundamentally suboptimal. This paper reveals that applying Reinforcement Learning (RL) directly to intermediate pre-training checkpoints (some as early as 4B tokens) can match the performance of the full standard pipeline. Crucially, RL applied early expands a model's reasoning capabilities, whereas RL applied after SFT merely sharpens them.

The "Invisible Leash" of SFT

In modern AI development, SFT (Supervised Fine-Tuning) is often seen as the necessary bridge between raw pre-training and RL. However, the authors argue that SFT acts as an "invisible leash." When RL follows SFT, it inherits a narrowed distribution. It gets better at picking the "best" answer (pass@1) but loses the diversity of thought required for complex problem-solving (pass@k).

By bypassing the initial SFT and going straight from a base checkpoint to RL, the model is forced to explore and discover reasoning paths on its own. The result? A model that doesn't just memorize solutions but actually expands its latent reasoning space.

Methodology: Reinventing the Pipeline

The researchers utilized a 1B parameter model (architecture based on OLMo2) and tested several training permutations:

  1. Direct RL: RL applied to intermediate base checkpoints.
  2. SFT: Standard supervision with ground-truth solutions.
  3. SFT→RL: The standard industrial pipeline.
  4. Parallel RL+SFT: A new hybrid approach averaging updates from both worlds.

The Architecture of the Experiment

Model Architecture and Pipeline Overview Figure 1: Comparison of different post-training recipes applied to intermediate pre-training checkpoints.

Key Insight 1: RL Works Surprisingly Early

Common wisdom suggests that a model must be "mature" (fully pre-trained) for RL to take hold. This study shatters that assumption. On the GSM8K benchmark, RL was highly effective at just 4B tokens—well before the "Chinchilla-optimal" point. By 10B tokens, the Direct RL model was already matching the performance of a model that had undergone the full SFT→RL treatment.

Key Insight 2: Data Composition > Model Scale

When tasks get harder (like the MATH benchmark), simply scaling the model from 1B to 4B parameters didn't make RL more effective. Instead, the "secret sauce" was Targeted Pre-training Data. Adding 10B tokens of math-heavy data during pre-training provided the necessary "latent capability" for RL to bootstrap from, narrowing the gap significantly.

Data vs Scale comparison Figure 2: Why data composition is a stronger lever than model size for RL effectiveness.

Key Insight 3: The "Parallel" Path Forward

Perhaps the most exciting contribution is the Parallel Averaging algorithm. Instead of doing SFT then RL, the authors suggest doing them at the same time. By calculating independent gradients for SFT and RL objectives and averaging them, the model achieves the highest pass@32 (diversity of solutions) while—critically—avoiding the "SFT Tax" (the degradation of general capabilities like Wikipedia knowledge).

Parallel Training Results Figure 3: Parallel Averaging outperforms the standard pipeline while preserving general capability.

Critical Analysis & Conclusion

This paper is a wake-up call for researchers who view RL as merely an "alignment" tool.

Limitations: The study focused on verifiable rewards (math/code). It remains to be seen if "early RL" is as effective for subjective tasks (creative writing, summarizing) where rewards are typically derived from noisy Reward Models (RM) rather than binary ground truths.

The Takeaway: If you are building reasoning models, don't wait for your pre-training to finish. Start your RL "excursions" early, focus heavily on your data mixture, and consider parallelizing your objectives to maintain the model's general intelligence. The future of LLM training is likely not a sequential pipeline, but a multi-objective, simultaneous optimization.

Find Similar Papers

Try Our Examples

  • Search for recent studies that integrate on-policy Reinforcement Learning objectives directly into the LLM pre-training phase rather than post-training.
  • Which paper first identified the 'distribution sharpening' effect in RLHF, and how do their experimental settings compare to the Direct RL approach used here?
  • Explore the performance impact of parallel gradient averaging across different training objectives (e.g., Contrastive Learning and Next Token Prediction) in foundation models.
Contents
RL Excursions: Shaking the Foundations of the LLM Training Pipeline
1. TL;DR
2. The "Invisible Leash" of SFT
3. Methodology: Reinventing the Pipeline
3.1. The Architecture of the Experiment
4. Key Insight 1: RL Works Surprisingly Early
5. Key Insight 2: Data Composition > Model Scale
6. Key Insight 3: The "Parallel" Path Forward
7. Critical Analysis & Conclusion