SPD: Steering the Internal "Capability Subspace" to Solve the Self-Distillation Bottleneck
Self-Policy Distillation via Capability-Selective Subspace Projection
Self-Policy Distillation (SPD) is a novel self-distillation framework that improves Large Language Models (LLMs) by training them on curated versions of their own outputs. Using a method called Capability-Selective Subspace Projection, it achieves SOTA results across code, math, and QA tasks without requiring any external reward models or verifiers.
TL;DR
Self-Policy Distillation (SPD) introduces a paradigm shift in how we improve LLMs using their own data. By projecting internal activations into a specialized "capability subspace" during the data generation phase, SPD creates a cleaner, more focused training corpus. It removes the need for expensive external verifiers while achieving double-digit improvements in reasoning and coding.
Background Positioning: This is a SOTA-level contribution to the field of on-policy self-alignment, moving away from external feedback (RLHF/RLEF) toward internal representation-level control.
The "Garbage In, Redundancy Out" Problem
The biggest headache in self-training is signal dilution. When an LLM generates data, it doesn't just produce a solution; it produces a "soup" of:
- Task Signal: The actual logic required to solve the problem.
- Stylistic Noise: Verbose greetings, redundant explanations, or formatting artifacts.
- Model Errors: Hallucinations and incorrect logic.
Previous methods like SSD (Simple Self-Distillation) try to fix this by truncation or simple filtering. However, they can't separate the "reasoning capability" from the "stylistic patterns" and often overfit on specific domains like code.
Methodology: The Geometry of Correctness
SPD's core intuition is that the model's internal Key (K) and Value (V) representations contain specific directions (subspaces) that are highly sensitive to "correctness."
Phase 1: Subspace Extraction
Instead of looking at the whole sentence, the authors define Correctness-aligned loss. They only calculate gradients for tokens that actually matter (the final number in a math problem or the logic in an assertion). By performing SVD (Singular Value Decomposition) on these gradients, they identify a low-rank matrix that represents the "Capability Subspace."
Phase 2: Generation with Projection Hooks
During data generation, SPD "hooks" into the model. It doesn't change the weights; instead, it projects every KV activation onto that extracted subspace.
- The Result: The model effectively filters its own thoughts. It stops "rambling" and focuses on the logic patterns that lead to correct answers.

Experiments: More Than Just Code
The authors tested SPD across five backbones (Qwen2.5, Llama-3.1, etc.) and three major domains. The results were consistently superior to "Plain Self-Retraining."
| Metric | Base Model | Simple Self-Distillation | SPD (Ours) |
|---|---|---|---|
| MBPP (Code) | 17.0% | 18.3% | 25.5% |
| GSM8K (Math) | 11.0% | 12.0% | 22.0% |
| BBH (QA) | 32.7% | 36.0% | 38.7% |
The "Style over Substance" Test
One of the most revealing parts of the paper is Figure 3. You can see that while the Base Model and SSD methods are verbose and include unnecessary print statements, the SPD-generated data is compact and implementation-focused. It’s not just "cleaner" data; it is "higher-density" intelligence.

Deep Insight: Out-of-Domain Generalization
Perhaps the most surprising finding is that a subspace extracted for Multiple-Choice QA can actually help the model perform better in Math and Code. This suggests that the "capability subspace" isn't just about a specific dataset; it captures a broader "reasoning policy" within the model's hidden layers.
Critical Analysis & Conclusion
Takeaway: SPD proves that LLMs already "know" how to be correct; their problem is that their correct reasoning is buried under noise. By using subspace projection as a lens, we can extract the gold from the gravel without needing a human or a stronger model to point it out.
Limitations: While powerful, the method currently requires a small "calibration set" (around 50-100 examples) to define the subspace. Future iterations might find ways to identify these subspaces completely unsupervised.
Future Work: This technique opens the door to "Steerable Self-Distillation," where one could theoretically isolate and amplify specific traits like creativity, conciseness, or safety purely through internal geometric projection.
