OPRD: Lifting LLM Distillation from Probabilities to Internal Representations

OPRD: On-Policy Representation Distillation

2026-06-01
Shenzhi Yang, Guangcheng Zhu, Bowen Song, Haobo Wang, Mingxuan Xia, Xing Zheng, Yingfan Ma, Zhongqi Chen, Weiqiang Wang, Gang Chen
Summary
Problem
Method
Results
Takeaways
Abstract

OPRD (On-Policy Representation Distillation) is a novel distillation framework for LLMs that shifts supervision from the output logit space to the intermediate hidden-state space. Applied to mathematical reasoning tasks, it enables student models to fully close the performance gap with teachers, achieving SOTA results on AIME and AIMO benchmarks.

TL;DR

On-Policy Representation Distillation (OPRD) is a breakthrough in LLM post-training. By supervising the "internal thoughts" (hidden states) of a model rather than just its "final words" (logits), OPRD eliminates the training noise that plagues current methods. It allows a 1.5B student model to effectively match a teacher's performance on elite math competitions like AIME, while training 44% faster and using half the memory.

The Problem: Why Logit-Matching Stagnates

Current On-Policy Distillation (OPD) methods, such as those used in training DeepSeek or Qwen, treat the teacher model as a "black-box probability oracle." They focus entirely on the output layer. This approach hits two walls:

  1. The Variance Trap: Estimating a KL-divergence over a 150k-token vocabulary using just one sampled token is inherently noisy. Late in training, this noise overwhelms the signal, causing model performance to plateau.
  2. The Information Bottleneck: The Language Model (LM) head is a "lossy" projection. It compresses high-dimensional hidden states () into vocabulary distributions. In this process, the structural "how" of the teacher's reasoning is lost.

Methodology: Seeing Through the Teacher's Eyes

OPRD moves the supervision target before the LM head. Instead of asking the student to predict the same next token as the teacher, OPRD demands the student replicate the teacher's internal hidden representations () across multiple layers.

Mathematical Intuition

The authors prove two critical theorems:

  • Theorem 1 (Zero Variance): Unlike policy gradients, the OPRD MSE loss is deterministic for a given rollout. This provides a stable signal even when the student is nearly as good as the teacher.
  • Theorem 2 (Breaking the Bottleneck): The LM head has a massive "effective null space." Many different internal hidden states can produce the same output distribution. OPRD constrains these hidden states directly, forcing the student to inhabit the same manifold as the teacher.

OPRD Architecture Figure 1: OPRD extracts supervision before the LM-head projection, exposing structural information that output-space objectives discard.

Experimental Results: Closing the Gap

The authors tested OPRD on grueling mathematical reasoning tasks (AIME 2024, AIME 2025, and AIMO).

  • Accuracy: OPRD reached 49.8 on AIME24, virtually tied with the teacher (50.8), while standard OPD variants plateaued significantly lower (~42-47).
  • Conciseness: Interestingly, OPRD-trained models produced shorter, more efficient reasoning chains (~5,700 tokens vs. ~7,000 for OPD), suggesting better reasoning "intuition."
  • Efficiency: Because the [Batch, Length, Vocab] tensor is never materialized for the loss, OPRD is a "training free lunch," saving up to 54% in peak activation memory.

Accuracy Comparison Figure 2: Training dynamics show OPRD (navy) climbing monotonically toward teacher performance (dashed line) while traditional OPD (green/blue) plateaus.

Deep Insights: The "Last-K" Strategy

A fascinating mechanistic discovery in the paper is that student hidden states diverge most from the teacher at the end of a response—where the final answer is committed. By focusing representation distillation on the last 2000 tokens of a Chain-of-Thought, OPRD achieves maximum impact with minimal compute.

Conclusion & Future Outlook

OPRD proves that the internal states of Transformers are a goldmine for distillation. While currently limited to "same-architecture" pairs (e.g., distilling a 1.5B model from another 1.5B model with different training history), it opens the door for:

  • Self-Distillation: Using privileged info (like ground truth) to guide the model's own hidden states.
  • Efficiency: Faster, more stable RL pipelines for reasoning models.

The future of LLM training may rely less on "what" the model says during training, and more on "how" it constructs its internal representation of the world.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize intermediate layer representations for on-policy reinforcement learning or distillation in Large Language Models.
  • Which study first identified the ill-conditioned singular spectrum of the LM-head as an information bottleneck in transformer models?
  • Explore methods for cross-architecture representation distillation that allow a smaller student to align with a larger teacher's hidden state space.
Contents
OPRD: Lifting LLM Distillation from Probabilities to Internal Representations
1. TL;DR
2. The Problem: Why Logit-Matching Stagnates
3. Methodology: Seeing Through the Teacher's Eyes
3.1. Mathematical Intuition
4. Experimental Results: Closing the Gap
5. Deep Insights: The "Last-K" Strategy
6. Conclusion & Future Outlook