Polar: Training Agentic RL on Any Harness Without "Opening the Box"

Polar: Agentic RL on Any Harness at Scale

2026-05-01
Binfeng Xu, Hao Zhang, Shaokun Zhang, Songyang Han, Mingjie Liu, Jian Hu, Shizhe Diao, Zhenghui Jin, Yunheng Zou, Michael Demoret, Jan Kautz, Yi Dong
Summary
Problem
Method
Results
Takeaways
Abstract

Polar is a scalable asynchronous Reinforcement Learning (RL) rollout framework that enables training of LLM agents using arbitrary, unmodified "as-is" harnesses. By proxying LLM API calls, Polar reconstructs token-faithful trajectories from black-box agent interactions, achieving significant gains on SWE-Bench Verified (e.g., +22.6 points for Qwen3.5-4B using the Codex harness).

TL;DR

Reinforcement Learning for agents is traditionally hard because adapting a complex coding or browsing harness into an RL environment is a systems nightmare. Polar solves this by moving the integration boundary to the LLM API itself. By proxying the agent's model calls, Polar captures trajectories and rewards from unmodified harnesses, enabling scalable RL across tools like Claude Code or Codex. Results show massive performance jumps (up to +22.6% on SWE-Bench) by simply training the model on the exact execution paths of these native harnesses.

The "Integration Tax" in Agentic RL

As agents move from simple chat to solving GitHub issues (SWE-Bench) or navigating OS environments, the "harness"—the software that lets the agent use tools and manage context—becomes incredibly complex.

Current SOTA RL frameworks like SkyRL or Agent Lightning often demand that you rewrite your agent to fit their specific Python API or SDK. This creates a "tax":

  1. Engineering Overhead: Manually porting a 10k-line coding harness into a Gym interface.
  2. Fidelity Loss: Rewriting the harness often changes how context is compacted or how tools are called, leading to a mismatch between training and deployment.
  3. Closed Systems: Some harnesses are distributed as binaries or closed-source tools where "opening the box" is impossible.

Methodology: The API Proxy as an RL Interface

Polar’s core insight is that while agent harnesses vary wildly, they all eventually call a model. Polar places a Model API Proxy between the harness and the inference server.

Polar Architecture Overview

1. Black-Box Rollouts

Polar runs the official agent harness (e.g., anthropic-code or codex) inside a container. It tricks the harness into sending requests to a local Polar gateway instead of OpenAI/Anthropic. The gateway:

  • Normalizes the request format.
  • Captures token IDs and logprobs directly from the local inference engine (avoiding retokenization drift).
  • Forwards the response back to the harness.

2. Token-Faithful Prefix Merging

Capturing every individual request is noisy. Polar uses a Prefix Merging strategy to reconstruct a single coherent multi-turn trajectory from many scattered API calls. It identifies "chains" where the prompt of turn contains the tokens of turn as a prefix. This maintains token fidelity: the trainer sees exactly what the model "felt" during generation.

Trajectory Reconstruction Example

Experiments: Massive Gains on Real-World Coding

The researchers tested Polar on SWE-Bench Verified, training a Qwen3.5-4B model using the GRPO (Group Relative Policy Optimization) algorithm.

The results are striking. When training on the Codex harness—which uses a specific tool schema the base Qwen model wasn't used to—performance jumped from 3.8% to 26.4%.

HarnessBase Model (4B)Polar RLImprovement
Codex3.8%26.4%+22.6
Claude Code29.8%34.6%+4.8
Pi Agent34.2%40.4%+6.2

These curves prove that RL can "teach" a model to speak the specific language of a harness, even if that harness was never intended for training.

SWE-Gym GRPO Training Curves

System Efficiency: Hiding the Cost of Setup

Agent rollouts are "slow" (minutes) compared to training steps (seconds). Polar uses Asynchronous Rollout Staging. While the GPU is busy running an agent, the CPU is already "pre-warming" the next session's Docker container and setting up the environment.

By using Prefix Merging, Polar also minimizes the number of samples sent to the GPU trainer, reducing wall-clock training time by over 5x compared to a naive per-request approach.

Critical Insight & Future Outlook

Polar shifts the philosophy of Agent RL from "Build-to-Train" to "Listen-to-Train." It recognizes that the engineering effort invested in product-level agent harnesses is too valuable to throw away during the training phase.

Limitations: Currently, Polar is most effective for "append-only" conversations. If a harness uses extreme context manipulation (jumping back and forth in history), prefix merging becomes harder.

Conclusion: For teams wanting to optimize their agents for specific tools or proprietary platforms, Polar provides the most frictionless path yet to scalable, high-fidelity Reinforcement Learning.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use proxy-based instrumentation or API interception for training Large Language Model agents instead of manual environment porting.
  • Which original research introduced the concept of "retokenization drift" in LLM agent training, and how does Polar's token-faithful reconstruction compare to its mitigation strategies?
  • Explore similar "rollout-as-a-service" architectures (like ProRL or SkyRL-Agent) and identify how they handle long-tail execution and CPU-heavy environment staging in RL pipelines.
Contents
Polar: Training Agentic RL on Any Harness Without "Opening the Box"
1. TL;DR
2. The "Integration Tax" in Agentic RL
3. Methodology: The API Proxy as an RL Interface
3.1. 1. Black-Box Rollouts
3.2. 2. Token-Faithful Prefix Merging
4. Experiments: Massive Gains on Real-World Coding
5. System Efficiency: Hiding the Cost of Setup
6. Critical Insight & Future Outlook