Polar: Training Agentic RL on Any Harness Without "Opening the Box"
Polar: Agentic RL on Any Harness at Scale
Polar is a scalable asynchronous Reinforcement Learning (RL) rollout framework that enables training of LLM agents using arbitrary, unmodified "as-is" harnesses. By proxying LLM API calls, Polar reconstructs token-faithful trajectories from black-box agent interactions, achieving significant gains on SWE-Bench Verified (e.g., +22.6 points for Qwen3.5-4B using the Codex harness).
TL;DR
Reinforcement Learning for agents is traditionally hard because adapting a complex coding or browsing harness into an RL environment is a systems nightmare. Polar solves this by moving the integration boundary to the LLM API itself. By proxying the agent's model calls, Polar captures trajectories and rewards from unmodified harnesses, enabling scalable RL across tools like Claude Code or Codex. Results show massive performance jumps (up to +22.6% on SWE-Bench) by simply training the model on the exact execution paths of these native harnesses.
The "Integration Tax" in Agentic RL
As agents move from simple chat to solving GitHub issues (SWE-Bench) or navigating OS environments, the "harness"—the software that lets the agent use tools and manage context—becomes incredibly complex.
Current SOTA RL frameworks like SkyRL or Agent Lightning often demand that you rewrite your agent to fit their specific Python API or SDK. This creates a "tax":
- Engineering Overhead: Manually porting a 10k-line coding harness into a Gym interface.
- Fidelity Loss: Rewriting the harness often changes how context is compacted or how tools are called, leading to a mismatch between training and deployment.
- Closed Systems: Some harnesses are distributed as binaries or closed-source tools where "opening the box" is impossible.
Methodology: The API Proxy as an RL Interface
Polar’s core insight is that while agent harnesses vary wildly, they all eventually call a model. Polar places a Model API Proxy between the harness and the inference server.

1. Black-Box Rollouts
Polar runs the official agent harness (e.g., anthropic-code or codex) inside a container. It tricks the harness into sending requests to a local Polar gateway instead of OpenAI/Anthropic. The gateway:
- Normalizes the request format.
- Captures token IDs and logprobs directly from the local inference engine (avoiding retokenization drift).
- Forwards the response back to the harness.
2. Token-Faithful Prefix Merging
Capturing every individual request is noisy. Polar uses a Prefix Merging strategy to reconstruct a single coherent multi-turn trajectory from many scattered API calls. It identifies "chains" where the prompt of turn contains the tokens of turn as a prefix. This maintains token fidelity: the trainer sees exactly what the model "felt" during generation.

Experiments: Massive Gains on Real-World Coding
The researchers tested Polar on SWE-Bench Verified, training a Qwen3.5-4B model using the GRPO (Group Relative Policy Optimization) algorithm.
The results are striking. When training on the Codex harness—which uses a specific tool schema the base Qwen model wasn't used to—performance jumped from 3.8% to 26.4%.
| Harness | Base Model (4B) | Polar RL | Improvement |
|---|---|---|---|
| Codex | 3.8% | 26.4% | +22.6 |
| Claude Code | 29.8% | 34.6% | +4.8 |
| Pi Agent | 34.2% | 40.4% | +6.2 |
These curves prove that RL can "teach" a model to speak the specific language of a harness, even if that harness was never intended for training.

System Efficiency: Hiding the Cost of Setup
Agent rollouts are "slow" (minutes) compared to training steps (seconds). Polar uses Asynchronous Rollout Staging. While the GPU is busy running an agent, the CPU is already "pre-warming" the next session's Docker container and setting up the environment.
By using Prefix Merging, Polar also minimizes the number of samples sent to the GPU trainer, reducing wall-clock training time by over 5x compared to a naive per-request approach.
Critical Insight & Future Outlook
Polar shifts the philosophy of Agent RL from "Build-to-Train" to "Listen-to-Train." It recognizes that the engineering effort invested in product-level agent harnesses is too valuable to throw away during the training phase.
Limitations: Currently, Polar is most effective for "append-only" conversations. If a harness uses extreme context manipulation (jumping back and forth in history), prefix merging becomes harder.
Conclusion: For teams wanting to optimize their agents for specific tools or proprietary platforms, Polar provides the most frictionless path yet to scalable, high-fidelity Reinforcement Learning.
