SPD: Steering the Internal "Capability Subspace" to Solve the Self-Distillation Bottleneck

Self-Policy Distillation via Capability-Selective Subspace Projection

2026-05-01
Guangya Hao, Yitong Shang, Yunbo Long, Zhuokai Zhao, Hanxue Liang
Summary
Problem
Method
Results
Takeaways
Abstract

Self-Policy Distillation (SPD) is a novel self-distillation framework that improves Large Language Models (LLMs) by training them on curated versions of their own outputs. Using a method called Capability-Selective Subspace Projection, it achieves SOTA results across code, math, and QA tasks without requiring any external reward models or verifiers.

TL;DR

Self-Policy Distillation (SPD) introduces a paradigm shift in how we improve LLMs using their own data. By projecting internal activations into a specialized "capability subspace" during the data generation phase, SPD creates a cleaner, more focused training corpus. It removes the need for expensive external verifiers while achieving double-digit improvements in reasoning and coding.

Background Positioning: This is a SOTA-level contribution to the field of on-policy self-alignment, moving away from external feedback (RLHF/RLEF) toward internal representation-level control.

The "Garbage In, Redundancy Out" Problem

The biggest headache in self-training is signal dilution. When an LLM generates data, it doesn't just produce a solution; it produces a "soup" of:

  1. Task Signal: The actual logic required to solve the problem.
  2. Stylistic Noise: Verbose greetings, redundant explanations, or formatting artifacts.
  3. Model Errors: Hallucinations and incorrect logic.

Previous methods like SSD (Simple Self-Distillation) try to fix this by truncation or simple filtering. However, they can't separate the "reasoning capability" from the "stylistic patterns" and often overfit on specific domains like code.

Methodology: The Geometry of Correctness

SPD's core intuition is that the model's internal Key (K) and Value (V) representations contain specific directions (subspaces) that are highly sensitive to "correctness."

Phase 1: Subspace Extraction

Instead of looking at the whole sentence, the authors define Correctness-aligned loss. They only calculate gradients for tokens that actually matter (the final number in a math problem or the logic in an assertion). By performing SVD (Singular Value Decomposition) on these gradients, they identify a low-rank matrix that represents the "Capability Subspace."

Phase 2: Generation with Projection Hooks

During data generation, SPD "hooks" into the model. It doesn't change the weights; instead, it projects every KV activation onto that extracted subspace.

  • The Result: The model effectively filters its own thoughts. It stops "rambling" and focuses on the logic patterns that lead to correct answers.

Overall Architecture of SPD

Experiments: More Than Just Code

The authors tested SPD across five backbones (Qwen2.5, Llama-3.1, etc.) and three major domains. The results were consistently superior to "Plain Self-Retraining."

MetricBase ModelSimple Self-DistillationSPD (Ours)
MBPP (Code)17.0%18.3%25.5%
GSM8K (Math)11.0%12.0%22.0%
BBH (QA)32.7%36.0%38.7%

The "Style over Substance" Test

One of the most revealing parts of the paper is Figure 3. You can see that while the Base Model and SSD methods are verbose and include unnecessary print statements, the SPD-generated data is compact and implementation-focused. It’s not just "cleaner" data; it is "higher-density" intelligence.

Self-Generated Data Comparison

Deep Insight: Out-of-Domain Generalization

Perhaps the most surprising finding is that a subspace extracted for Multiple-Choice QA can actually help the model perform better in Math and Code. This suggests that the "capability subspace" isn't just about a specific dataset; it captures a broader "reasoning policy" within the model's hidden layers.

Critical Analysis & Conclusion

Takeaway: SPD proves that LLMs already "know" how to be correct; their problem is that their correct reasoning is buried under noise. By using subspace projection as a lens, we can extract the gold from the gravel without needing a human or a stronger model to point it out.

Limitations: While powerful, the method currently requires a small "calibration set" (around 50-100 examples) to define the subspace. Future iterations might find ways to identify these subspaces completely unsupervised.

Future Work: This technique opens the door to "Steerable Self-Distillation," where one could theoretically isolate and amplify specific traits like creativity, conciseness, or safety purely through internal geometric projection.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Singular Value Decomposition (SVD) on gradients to steer Large Language Model activations or fine-tuning.
  • Which paper first introduced the concept of "Representation Engineering" for LLM steering, and how does Self-Policy Distillation evolve this concept?
  • Find studies investigating "out-of-domain transfer in self-distillation" to see if steering in one capability (e.g., coding) consistently improves others (e.g., logic).
Contents
SPD: Steering the Internal "Capability Subspace" to Solve the Self-Distillation Bottleneck
1. TL;DR
2. The "Garbage In, Redundancy Out" Problem
3. Methodology: The Geometry of Correctness
3.1. Phase 1: Subspace Extraction
3.2. Phase 2: Generation with Projection Hooks
4. Experiments: More Than Just Code
4.1. The "Style over Substance" Test
5. Deep Insight: Out-of-Domain Generalization
6. Critical Analysis & Conclusion