DistIL: Fixing Credit Assignment in LLM Reasoning with Distributional DAgger

4

Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces DistIL, a distributional variant of the DAgger algorithm designed for reinforcement learning from rich feedback (e.g., execution traces, logs, or expert corrections). It optimizes a forward cross-entropy objective with future-aware credit assignment, outperforming SOTA RLVR and self-distillation baselines like SDPO and GRPO across reasoning, coding, and mathematics.

Executive Summary

Reinforcement Learning from Verifiable Rewards (RLVR) is the current engine behind reasoning models like DeepSeek-R1 and OpenAI's o1. However, the standard recipe is surprisingly crude: it samples responses and gives a binary "1" or "0" reward. This paper argues that this fails credit assignment—the model doesn't know which step led to the right answer. The authors propose DistIL, a new framework that uses "rich feedback" (like error logs and traces) via a distributional version of the classic DAgger algorithm. By using a forward cross-entropy objective and a "future-aware" gradient, DistIL achieves monotonic policy improvement and SOTA results in science, coding, and math.

The Problem: The "Sparse Reward" Trap

Current methods like GRPO or SDPO suffer from two fundamental flaws:

  1. Divergence Mismatch: Many self-distillation methods use reverse-KL. The authors prove that even if your "teacher" is better than your "student," a reverse-KL update can actually decrease the student's reward (see Proposition 2). This happens because reverse-KL is "mode-seeking" and can over-penalize moderate but good actions.
  2. Local Myopia: Most methods use "local" gradients, updating tokens based only on immediate mismatch. This ignores how an early decision (e.g., step 1 of a math proof) affects the state distribution 50 tokens later.

Methodology: The DistIL Approach

DistIL (Distributional Imitation Learning) moves away from the simple binary reward and focuses on a feedback-conditioned Teacher.

1. Forward Cross-Entropy

Unlike reverse-KL, the forward cross-entropy objective ensures that the student always moves toward the teacher's distribution. The paper provides a mathematical guarantee of monotonic policy improvement: if the teacher is better, the update will improve the student.

2. Future-Aware Credit Assignment

This is the core "How." Instead of just looking at the token at time , DistIL's gradient (Equation 5) includes a future-credit term. It calculates how much the student and teacher will disagree in the future and propagates that signal back to the current token.

Overall Performance Comparison Figure 1: DistIL vs SDPO. Note the stability and monotonic growth of DistIL across scientific domains compared to the oscillations of SDPO.

Experimental Battleground

The researchers tested DistIL against heavyweights like GRPO and SDPO:

  • Science (SciKnowEval): DistIL showed massive gains in Physics (+9.6 points) and Chemistry.
  • Coding (LiveCodeBench): DistIL exploited execution logs (errors, unit tests) better than RLVR, which can't "read" an error log.
  • Hard Math (OmniMath): In "zero-pass" regimes—where the base model is too weak to ever find the right answer—GRPO fails (it gets no reward). DistIL, by imitating a teacher conditioned on ground-truth solutions, successfully "bootstraps" the model.

Coding Performance on LCBv6 Figure 2: Coding results. DistIL consistently leads in both Best@k and Majority@k metrics.

Deep Insight: Pass@N and Success Likelihood

One of the most profound theoretical contributions (Proposition 5) is that DistIL's objective is a lower bound on the likelihood of success. This explains why it excels at Pass@N metrics. By minimizing cross-entropy with a successful teacher, the student is implicitly maximizing its own probability of finding at least one correct path.

Critical Analysis & Conclusion

DistIL proves that the "richness" of feedback matters as much as the amount of data. By shifting from outcome-based RL to distributional imitation with sequence-level awareness, we solve the two biggest headaches in reasoning models: credit assignment and training instability.

Limitations: The method relies on having a "better" teacher. If the feedback is noisy or the teacher-student gap is too large, the "concentrability" (overlap) might suffer. However, as an on-policy method, it naturally mitigates distribution shift better than traditional Behavior Cloning.

Future Work: The logical next step is applying DistIL to multi-modal reasoning (CV/Audio) where "verifiable rewards" are even harder to come by, but "rich feedback" (like descriptions) is abundant.

Find Similar Papers

Try Our Examples

  • Find recent papers that compare forward cross-entropy versus reverse-KL divergence in the context of on-policy LLM self-distillation.
  • Which studies first identified the 'local credit assignment' failure in DAgger-like algorithms for large language models?
  • Explore if DistIL's future-aware credit assignment can be combined with Process-Reward Models (PRMs) for even denser supervision in non-verifiable domains.
Contents
DistIL: Fixing Credit Assignment in LLM Reasoning with Distributional DAgger
1. Executive Summary
2. The Problem: The "Sparse Reward" Trap
3. Methodology: The DistIL Approach
3.1. 1. Forward Cross-Entropy
3.2. 2. Future-Aware Credit Assignment
4. Experimental Battleground
5. Deep Insight: Pass@N and Success Likelihood
6. Critical Analysis & Conclusion