SDAR: Taming Multi-Turn Agent Trajectories with Gated Self-Distillation

Self-Distilled Agentic Reinforcement Learning

2026-01-01
Zhengxi Lu, Zhiyuan Yao, Zhuowen Han, Zi-Han Wang, Jinyang Wu, Qi Gu, Xunliang Cai, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
Summary
Problem
Method
Results
Takeaways
Abstract

SDAR (Self-Distilled Agentic Reinforcement Learning) is a post-training framework for multi-turn LLM agents that integrates Reinforcement Learning (GRPO) with a gated On-Policy Self-Distillation (OPSD) objective. It utilizes a sigmoid-based token-level gating mechanism to selectively distill dense guidance from a privileged teacher, achieving SOTA results across ALFWorld, WebShop, and Search-QA benchmarks.

TL;DR

Training LLM agents to handle long-horizon, multi-turn tasks is notoriously difficult due to "sparse rewards" and "compounding errors." While Reinforcement Learning (RL) provides the target, it lacks the resolution for token-by-token guidance. SDAR (Self-Distilled Agentic Reinforcement Learning) bridges this gap by combining RL with Gated On-Policy Self-Distillation. It uses a "privileged teacher" (the model itself plus expert skills) to provide dense feedback, but crucially uses a Sigmoid Gate to ignore the teacher when it becomes unreliable or noisy.

The Motivation: Why Naive Distillation Fails in Multi-Turn

When an agent interacts with an environment over multiple turns, two major issues arise:

  1. Multi-turn Instability: Once a student makes a mistake and drifts from the teacher's trajectory, the teacher's once-helpful advice becomes irrelevant or confusing, leading to a surge in KL divergence and performance collapse.
  2. Asymmetric Trust: In Self-Distillation, the "teacher" is just the "student" plus some extra context (like retrieved skills). If the skills are poor, the teacher's rejection of a student's token might actually be wrong.

The authors observed that in over 50% of tokens, the "privileged teacher" actually assigned a lower probability than the student. If we distilled everything uniformly, we would be training the model to be worse.

Methodology: The Gated Auxiliary Objective

SDAR maintains the stability of GRPO (Group Relative Policy Optimization) as the backbone but adds a selective distillation loss.

1. The Sigmoid Gate ()

Instead of uniform distillation, SDAR calculates a "Gap" (): This gap is passed through a sigmoid function to create a gate value :

  • Positive Gap: Teacher likes the token more than the student Gate opens Strong Distillation.
  • Negative Gap: Teacher is confused or skills are irrelevant Gate closes Feedback is attenuated.

2. Framework Overview

SDAR Framework Architecture The framework uses a dual-pathway approach: the policy is optimized via environment rewards (Verifier) and the gated token-level signals from the Teacher branch.

Experimental Results: Internalizing Knowledge

SDAR was tested across three major benchmarks: ALFWorld (embodied tasks), WebShop (shopping agents), and Search-QA (retrieval-augmented reasoning).

ModelMethodALFWorld (Avg)WebShop Acc
Qwen2.5-7BGRPO (Baseline)72.6%37.6%
Qwen2.5-7BSDAR (Ours)82.8% (+10.2)73.0% (+35.4)

Key Insight: Better than "Cheating"

One of the most impressive results is Knowledge Internalization. Many agents perform well as long as you give them the "cheat sheet" (skills in the prompt). However, SDAR-trained models perform better without the skills at inference time than other models do with them.

Comparison of Training Dynamics

The graph above shows that while naive GRPO+OPSD often fluctuates and fails, SDAR provides a steady climb in success rate, proving that gated distillation prevents the gradients from "overwhelming" the RL signal.

Critical Analysis & Conclusion

Takeaway: SDAR proves that dense signal matters, but selective dense signal is the secret sauce for agents. By treating OPSD as a "gated" auxiliary task, the model learns the logic of the skills without becoming dependent on their presence in the prompt.

Limitations: The performance still relies on the base model's ability to at least sometimes generate high-quality teacher signals with privileged context. If the "privileged context" provides zero advantage, the distillation gate will stay closed, and the model reverts to standard RL.

Future Outlook: This gating strategy could be a paradigm shift for Self-Evolving Agents. In the future, agents might use this to decide exactly which parts of a self-reflection or a retrieved document are worth "memorizing" into their weights via RL fine-tuning.

Find Similar Papers

Try Our Examples

  • Search for recent papers that address the "compounding error" or "drift" problem specifically in multi-turn LLM agents using Reinforcement Learning.
  • Which original research introduced the concept of "privileged context" in the training of autonomous agents, and how does SDAR's gated approach build upon that foundation?
  • Are there studies applying similar token-level gating or confidence-based reweighting strategies to Multi-modal LLM agents or robotic control tasks?
Contents
SDAR: Taming Multi-Turn Agent Trajectories with Gated Self-Distillation
1. TL;DR
2. The Motivation: Why Naive Distillation Fails in Multi-Turn
3. Methodology: The Gated Auxiliary Objective
3.1. 1. The Sigmoid Gate ($\Delta_t$)
3.2. 2. Framework Overview
4. Experimental Results: Internalizing Knowledge
4.1. Key Insight: Better than "Cheating"
5. Critical Analysis & Conclusion