ECHO: Turning Terminal Feedback into a "World Model" for CLI Agents

ECHO: Terminal Agents Learn World Models for Free

2026-05-01
Vaishnavi Shrivastava, Piero Kauffmann, Ahmed Awadallah, Dimitris Papailiopoulos
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces ECHO (Environment Cross-entropy Hybrid Objective), a training method for CLI agents that supplements standard Reinforcement Learning (GRPO) with a self-supervised task: predicting terminal outputs (stdout, errors, logs). By treating environment responses as dense supervision signals, ECHO achieves SOTA results on TerminalBench-2.0, nearly doubling the pass@1 rate for Qwen3-8B and 14B models.

TL;DR

CLI agents typically learn through trial and error, but traditional Reinforcement Learning (RL) often "throws away" the most informative part of the trial: the environment's response. ECHO (Environment Cross-entropy Hybrid Objective) changes this by forcing the model to predict what the terminal will say next. This simple auxiliary loss doubles performance on TerminalBench-2.0 and allows models to match expert-level performance without a single expert demonstration.

Motivation: Stop Wasting the Step-by-Step Signals

In a standard RL setup for agents (like using GRPO), the model gets a reward only at the very end of a task. If the agent fails—which happens in 85% of cases for mid-sized models—the entire rollout provides almost zero policy-gradient signal.

However, the stdout, logs, and error messages generated during those failures are gold mines of information. If an agent runs ls and sees a file list, it has learned something about the state of the world. The authors of ECHO argue that "good prediction implies good understanding." By training the agent to predict these environment tokens, we are effectively training it to build a mental World Model of the terminal.

Methodology: The ECHO Objective

The beauty of ECHO lies in its simplicity. It requires no extra rollouts and no teacher models. It leverages the fact that environment observations are already in the context window.

The Hybrid Loss Function

ECHO modifies the standard training objective into a dual-task problem:

  1. Action Loss (GRPO): Optimize action tokens based on sparse rewards (Standard RL).
  2. Environment Loss (ECHO): Minimize cross-entropy on observation tokens (Self-Supervision).

ECHO Architecture and Concept Figure 1: ECHO turns terminal feedback into dense supervision during agent RL.

By using a single forward pass, the model updates its weights to not only pick winning actions but also to "understand" why the terminal responds the way it does.

Critical Insight: Learning Terminal Dynamics

Does ECHO actually learn how a terminal works? The authors tested this by measuring how well the model could predict outputs on trajectories it didn't generate (off-policy).

While standard GRPO models showed almost no improvement in predicting environment behavior, ECHO-trained models saw a sharp drop in cross-entropy. This proves the model isn't just memorizing its own mistakes—it is learning the underlying physics of the CLI environment.

Performance Curves Figure 2: Pass-rate training curves showing ECHO (Pink) consistently outperforming the GRPO baseline (Teal) across different benchmarks.

Key Results

  • Performance Leap: Qwen3-14B improved from a 5.17% pass rate to 10.79% on the rigorous TerminalBench-2.0.
  • Efficiency: ECHO-trained models reached peak performance 1.5x to 2.3x faster than those using standard GRPO.
  • Replacing Experts: On internal evaluations, ECHO applied to a base Qwen3-8B model matched the performance of a model that had been fine-tuned on 15,000 expert demonstrations. This suggests that self-supervised interaction can replace the need for expensive human/expert data.
  • Verifier-Free Adaptation: Remarkably, in some settings, the model could improve its task success just by "predicting the environment," even when the reward signal (verifier) was turned off.

Critical Analysis & Conclusion

Takeaway

ECHO proves that for agents, observations are supervision. The "World Model" doesn't need to be a separate component or a complex latent space; it can be learned directly via next-token prediction on the environment's literal responses.

Limitations

The effectiveness of ECHO depends on the "action-linked" nature of the feedback. In domains where terminal output is cryptic or disconnected from the agent's actions (e.g., complex orchestration), the benefit narrows. Furthermore, while it helps with the "interaction prior," it doesn't entirely replace the "strategy prior" that expert demonstrations provide for high-level planning.

Future Outlook

This work opens the door for a new era of autonomous agentic improvement. If an agent can learn simply by interacting with its environment and predicting the consequences, we may soon see agents that continuously refine their skills "in the wild" without needing constant human feedback.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize auxiliary prediction tasks or self-supervised world modeling to improve the sample efficiency of large language model agents in interactive environments.
  • Which paper first introduced the GRPO (Group-Relative Policy Optimization) algorithm, and how does ECHO specifically modify the loss masking compared to the original implementation?
  • Explore research that applies environment-token prediction or "learning from consequences" to other embodied AI domains like web navigation agents or robotic control in text-based environments.
Contents
ECHO: Turning Terminal Feedback into a "World Model" for CLI Agents
1. TL;DR
2. Motivation: Stop Wasting the Step-by-Step Signals
3. Methodology: The ECHO Objective
3.1. The Hybrid Loss Function
4. Critical Insight: Learning Terminal Dynamics
5. Key Results
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Outlook