DITTO: Mastering Social Intelligence via Verbal Reinforcement Learning

Reinforcing Human Behavior Simulation via Verbal Feedback Reinforcing Human Behavior Simulation via Verbal Feedback

2026-05-19
Weiwei Sun, Xuhui Zhou, Jiarui Liu, Weihua Du, Haojia Sun, Yiqing Xie, Qianou Ma, Sihao Chen, Mengting Wan, Longqi Yang, Pei Zhou, Sherry Wu, Sean Welleck, Graham Neubig, Yiming Yang, Maarten Sap
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces DITTO, a reinforcement learning (RL) framework for human behavior simulation that treats subjective verbal feedback as a primary training signal. By leveraging a new unified benchmark named SOUL (Simulation gym Of hUman-Like behavior) containing 10 diverse tasks, DITTO achieves a 36% average improvement over base models and outperforms GPT-4o on 6 out of 10 benchmarks.

TL;DR

Researchers from CMU and Microsoft have launched DITTO, a model designed to bridge the "Sim2Real" gap in human behavior simulation. Unlike standard RL that relies on abstract numerical scores, DITTO learns from descriptive verbal feedback (e.g., "you sounding too adversarial"). By training on the SOUL benchmark—a new suite of 10 social tasks—an 8B model was able to outperform much larger proprietary models by internalizing linguistic critiques during training.

The Problem: The "Scalar" Poverty of Social Learning

When training an LLM to solve a math problem, a binary "Correct/Incorrect" reward works perfectly. However, human behavior isn't binary. If a simulated patient is "too compliant" or a simulated student "lacks realistic misconceptions," a scalar reward of 0.6 provides zero guidance on what was wrong.

Existing simulators often suffer from:

  • Homogeneity: All simulated users sound the same.
  • Superhuman Bias: Models are too helpful or too logical to be human.
  • Feedback Loss: Multi-dimensional social nuances (trust, politeness, goal-alignment) are compressed into a single, uninformative number.

Methodology: Verbal Feedback as a First-Class Signal

DITTO (Reinforced Human Behavior Simulation via Verbal Feedback) changes the optimization objective. The core idea is to treat verbal critique as "privileged information"—available during training to guide the model, but not needed at inference.

The DITTO Loop:

  1. Draft Rollout (): The model acting as a specific persona generates a response.
  2. Verbal Critique (): An LLM judge provides a detailed text reflection on dimensions like Believability, Social Norms, and Goal Achievement.
  3. Refined Rollout (): The same model generates a new response, this time conditioned on the critique ().
  4. Joint Optimization: Using GRPO, both the "raw" attempt and the "refined" attempt are optimized. The model effectively learns to "distill" the wisdom of the critique into its base weights.

DITTO Architecture Overview

SOUL: A Unified Gym for Human-Like Behavior

To test DITTO, the authors introduced SOUL (Simulation gym Of hUman-Like behavior). It consolidates 10 tasks across 6 critical categories:

  • Theory of Mind (ToM): Understanding "who knows what" in complex scenes.
  • Character Role Play: Simulating literary figures with high fidelity.
  • Social Skill: Negotiating and interacting in multi-agent environments (Sotopia).
  • User/Learner/Persona Simulation: Acting as diverse proxies for clinical, educational, or social testing.

Experimental Breakthroughs

The results show that DITTO 8B is a "giant slayer." It achieved a 36% improvement over its base version and beat GPT-4o in 6 out of 10 tasks, particularly in high-stakes social interactions.

SOTopai Results Comparison

Why Verbal Feedback Wins over Scalar RL:

  • Learning Speed: Verbal feedback identifies the "failure mode" immediately (e.g., "you are being too dismissive"), whereas scalar RL requires the model to "guess" which part of its long output was the problem.
  • Safety & Secret-Keeping: In social simulations, models often leak secrets. DITTO significantly outperformed standard GRPO in preserving private information because the feedback could explicitly flag "you leaked the secret in turn 3."
  • Teacher-Student Gap: Analysis showed that the feedback-conditioned "Teacher" consistently outperformed the "Student," providing a stable and informative moving target for the policy to chase.

Deep Insight: Beyond Reasoning

DITTO represents a shift in LLM alignment. While most current research (like DeepSeek-R1) focuses on verifiable reasoning (math/code), DITTO proves that the same RL principles—specifically GRPO and self-reflection—can be applied to the "fuzzy" and subjective world of social intelligence.

Limitations & Future

While powerful, the framework currently relies on expensive LLM-as-a-judge pipelines. Future work will likely look into making these "critique sessions" more efficient and exploring how these social "world models" can help LLMs navigate real-world human-AI collaboration more safely.

Takeaway: If you want a model to act like a human, stop giving it grades and start giving it advice.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Group Relative Policy Optimization (GRPO) for tasks beyond mathematical reasoning or code generation.
  • Which study first introduced the concept of 'Reinforcement Learning from Text Feedback' (RLTF) and how does its use of Advantage-Weighted Regression compare to current policy gradient methods?
  • What are the current state-of-the-art benchmarks for 'Theory of Mind' (ToM) in LLMs that specifically focus on multi-party conversation and information asymmetry?
Contents
DITTO: Mastering Social Intelligence via Verbal Reinforcement Learning
1. TL;DR
2. The Problem: The "Scalar" Poverty of Social Learning
3. Methodology: Verbal Feedback as a First-Class Signal
3.1. The DITTO Loop:
4. SOUL: A Unified Gym for Human-Like Behavior
5. Experimental Breakthroughs
5.1. Why Verbal Feedback Wins over Scalar RL:
6. Deep Insight: Beyond Reasoning
6.1. Limitations & Future