Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
Summary
Problem
Method
Results
Takeaways
Abstract
This paper provides a theoretical framework using in-context weight prediction tasks to analyze the synergistic effects between LLM pre-training and post-training (SFT and RL). It reveals that while SFT excels with small, high-quality, "hard" datasets, Outcome Supervision/RL requires massive data scale to navigate a high-curvature loss landscape.
-
