Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models

Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models

Summary
Problem
Method
Results
Takeaways
Abstract

This paper provides a theoretical framework using in-context weight prediction tasks to analyze the synergistic effects between LLM pre-training and post-training (SFT and RL). It reveals that while SFT excels with small, high-quality, "hard" datasets, Outcome Supervision/RL requires massive data scale to navigate a high-curvature loss landscape.

-

Find Similar Papers

Try Our Examples

  • Search for recent studies on the "double descent" or "U-shaped" performance curves specifically within the Supervised Fine-Tuning (SFT) phase of Large Language Models.
  • What are the foundational papers discussing the "spectral radius" of Transformer transition matrices and how it relates to overthinking or stability in iterative Chain-of-Thought?
  • How does the data quality vs. scale trade-off analyzed in this paper compare to empirical findings in the training of DeepSeek-R1 or OpenAI o1 models?