HPT: Unifying SFT and RL for the Next Generation of LLM Post-Training

Towards a Unified View of Large Language Model Post-Training

2025-01-01
Xingtai Lv, Yuxin Zuo, Youbang Sun, Hongyi Liu, Yuntian Wei, Zhekai Chen, Lixuan He, Xuekai Zhu, Kaiyan Zhang, Bingning Wang, Ning Ding, Bowen Zhou, Bowen Zhou
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the Unified Policy Gradient Estimator (UPGE), a theoretical framework that unifies Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) as instances of a single optimization process. Leveraging this, the authors propose Hybrid Post-Training (HPT), which dynamically balances online exploration and offline exploitation, achieving SOTA results on benchmarks like AIME 2024 (+7 points over baselines).

TL;DR

LLM training is often viewed as a two-act play: first Supervised Fine-Tuning (SFT) to learn the format, then Reinforcement Learning (RL) to master reasoning. A new paper from Tsinghua University and Shanghai AI Lab proves these aren't two different acts, but the same math. By introducing the Unified Policy Gradient Estimator (UPGE), the authors reveal that SFT and RL are just different ends of a bias-variance spectrum. Their resulting algorithm, Hybrid Post-Training (HPT), dynamically switches between these signals, breaking the performance ceiling of traditional "SFT-then-RL" pipelines.

The Myth of Contradicting Objectives

In the traditional view:

  • SFT excels at exploitation: it forces the model to mimic high-quality offline demonstrations.
  • RL excels at exploration: it lets the model find its own path to a correct answer.

The problem? Doing them sequentially is expensive and brittle. If the SFT stage is too weak, RL exploration fails (the "Zero-RL" cold start problem). If RL is too aggressive, the model suffers from "catastrophic forgetting" of the reasoning patterns learned in SFT.

The authors propose a Common Objective: maximize expected success while maintaining proximity to a behavior policy. They prove that the gradients of SFT, PPO, GRPO, and even offline RL methods like SRFT can all be expressed as:

abla \pi_{ heta}$$ ## The Unified Policy Gradient Estimator (UPGE) This unification breaks the post-training process into four modular components as shown in the architecture below: ![Overall Architecture of UPGE](https://cdn.atominnolab.com/wisdoc/images/20260525-282b7249-f630-4946-9107-d9fdbe5726d1/page_000_block_010.png) 1. **Stabilization Mask ($\mathbb{1}_{stable}$)**: Like PPO clipping, it prevents wild distribution shifts. 2. **Reference Policy ($\pi_{ref}$)**: The denominator that reweights updates based on how common or rare a token was during sampling. 3. **Advantage Estimate ($\hat{A}$)**: The signal of "goodness." In SFT, this is implicitly +1 for demonstrations; in RL, it’s the reward relative to a baseline. 4. **Likelihood Gradient ($ abla \pi_{ heta}$)**: The engine that maps signals back to model weights. ### The Breakthrough: Hybrid Post-Training (HPT) Rather than a fixed sequence, HPT uses **Performance Feedback**. For every question: - If the model **fails** to solve the problem through rollouts ($P \leq \gamma$), HPT triggers an **SFT update** using high-quality offline data. - If the model **succeeds** ($P > \gamma$), it triggers an **RL update** to reinforce its own successful discovery. This creates a "just-in-time" learning system where the model only uses the "crutch" of SFT when it is truly lost. ## Experimental Results: Breaking the Boundary The effectiveness of HPT is most visible in complex reasoning tasks. On **AIME 2024**, HPT achieved a staggering **33.0%**, compared to just **19.4%** for pure GRPO. ![Performance Comparison on Reasoning Benchmarks](https://cdn.atominnolab.com/wisdoc/images/20260525-282b7249-f630-4946-9107-d9fdbe5726d1/page_012_block_002.png) ### Key Insights from the Data: - **Preserved Exploration**: Unlike pure SFT which can narrow a model's output variance, HPT maintains high **Pass@1024** scores. It expands the "capability boundary" rather than just making the model more confident in a narrow range. - **Efficiency**: HPT reduces training costs because it skips expensive rollouts on samples where it decides to perform SFT. - **Internalization**: Analysis shows that models trained with HPT actually internalize long-form reasoning patterns. Even when the ratio of SFT data drops late in training, the model's response length and reasoning depth do not regress. ## Conclusion and Future Outlook The "unified view" presented here simplifies the complex landscape of LLM alignment. By treating SFT and RL as a single optimization process, we move toward models that can autonomously bridge the gap between human instruction and self-discovered logic. **Limitations**: Currently, HPT relies on the existence of a binary verifier (rule-based rewards), making it ideal for Math and Code. Extending this to open-ended creative writing or subjective alignment will require integrating learned Reward Models into the UPGE framework. **Takeaway**: The future of post-training is not a pipeline, but a dynamic, unified policy optimization.

Find Similar Papers

Try Our Examples

  • Search for recent papers that explore "Online SFT" or dynamic data mixing strategies in LLM post-training beyond fixed ratios.
  • Which theoretical works originally established the connection between behavioral cloning and policy gradients in the context of KL-regularized RL?
  • Analyze research applying Hybrid Post-Training or similar performance-feedback mechanisms to non-mathematical domains like code generation or multimodal reasoning.
Contents
HPT: Unifying SFT and RL for the Next Generation of LLM Post-Training
1. TL;DR
2. The Myth of Contradicting Objectives
3. The Unified Policy Gradient Estimator (UPGE)
3.1. The Breakthrough: Hybrid Post-Training (HPT)
4. Experimental Results: Breaking the Boundary
4.1. Key Insights from the Data:
5. Conclusion and Future Outlook