SDAR: Taming Multi-Turn Agent Trajectories with Gated Self-Distillation
Self-Distilled Agentic Reinforcement Learning
SDAR (Self-Distilled Agentic Reinforcement Learning) is a post-training framework for multi-turn LLM agents that integrates Reinforcement Learning (GRPO) with a gated On-Policy Self-Distillation (OPSD) objective. It utilizes a sigmoid-based token-level gating mechanism to selectively distill dense guidance from a privileged teacher, achieving SOTA results across ALFWorld, WebShop, and Search-QA benchmarks.
TL;DR
Training LLM agents to handle long-horizon, multi-turn tasks is notoriously difficult due to "sparse rewards" and "compounding errors." While Reinforcement Learning (RL) provides the target, it lacks the resolution for token-by-token guidance. SDAR (Self-Distilled Agentic Reinforcement Learning) bridges this gap by combining RL with Gated On-Policy Self-Distillation. It uses a "privileged teacher" (the model itself plus expert skills) to provide dense feedback, but crucially uses a Sigmoid Gate to ignore the teacher when it becomes unreliable or noisy.
The Motivation: Why Naive Distillation Fails in Multi-Turn
When an agent interacts with an environment over multiple turns, two major issues arise:
- Multi-turn Instability: Once a student makes a mistake and drifts from the teacher's trajectory, the teacher's once-helpful advice becomes irrelevant or confusing, leading to a surge in KL divergence and performance collapse.
- Asymmetric Trust: In Self-Distillation, the "teacher" is just the "student" plus some extra context (like retrieved skills). If the skills are poor, the teacher's rejection of a student's token might actually be wrong.
The authors observed that in over 50% of tokens, the "privileged teacher" actually assigned a lower probability than the student. If we distilled everything uniformly, we would be training the model to be worse.
Methodology: The Gated Auxiliary Objective
SDAR maintains the stability of GRPO (Group Relative Policy Optimization) as the backbone but adds a selective distillation loss.
1. The Sigmoid Gate ()
Instead of uniform distillation, SDAR calculates a "Gap" (): This gap is passed through a sigmoid function to create a gate value :
- Positive Gap: Teacher likes the token more than the student Gate opens Strong Distillation.
- Negative Gap: Teacher is confused or skills are irrelevant Gate closes Feedback is attenuated.
2. Framework Overview
The framework uses a dual-pathway approach: the policy is optimized via environment rewards (Verifier) and the gated token-level signals from the Teacher branch.
Experimental Results: Internalizing Knowledge
SDAR was tested across three major benchmarks: ALFWorld (embodied tasks), WebShop (shopping agents), and Search-QA (retrieval-augmented reasoning).
| Model | Method | ALFWorld (Avg) | WebShop Acc |
|---|---|---|---|
| Qwen2.5-7B | GRPO (Baseline) | 72.6% | 37.6% |
| Qwen2.5-7B | SDAR (Ours) | 82.8% (+10.2) | 73.0% (+35.4) |
Key Insight: Better than "Cheating"
One of the most impressive results is Knowledge Internalization. Many agents perform well as long as you give them the "cheat sheet" (skills in the prompt). However, SDAR-trained models perform better without the skills at inference time than other models do with them.

The graph above shows that while naive GRPO+OPSD often fluctuates and fails, SDAR provides a steady climb in success rate, proving that gated distillation prevents the gradients from "overwhelming" the RL signal.
Critical Analysis & Conclusion
Takeaway: SDAR proves that dense signal matters, but selective dense signal is the secret sauce for agents. By treating OPSD as a "gated" auxiliary task, the model learns the logic of the skills without becoming dependent on their presence in the prompt.
Limitations: The performance still relies on the base model's ability to at least sometimes generate high-quality teacher signals with privileged context. If the "privileged context" provides zero advantage, the distillation gate will stay closed, and the model reverts to standard RL.
Future Outlook: This gating strategy could be a paradigm shift for Self-Evolving Agents. In the future, agents might use this to decide exactly which parts of a self-reflection or a retrieved document are worth "memorizing" into their weights via RL fine-tuning.
