Extrapolative Weight Averaging: Beyond the RL Frontier in Code Generation

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL

2026-05-01
Kunhao Zheng, Pierre Chambon, Juliette Decugis, Jonas Gehring, Taco Cohen, Benjamin Negrevergne, Gabriel Synnaeve
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Extrapolative Weight Averaging to navigate the correctness-efficiency frontier in code Reinforcement Learning (RL). By training on nested unit-test coverage, it achieves a 3.3% improvement in pass@250 on competitive programming benchmarks like LiveCodeBench (LCB)/hard.

TL;DR

Meta Researchers have discovered that Reinforcement Learning (RL) for code doesn't just "improve" models—it rotates them along a correctness-efficiency frontier. By using weight extrapolation between checkpoints trained on different test-case sizes, they can generate new, diverse policies that solve "impossible" problems without any extra training.

The Problem: The Hidden Trade-off in Code RL

In competitive programming, a solution must be two things: correct (logic) and efficient (complexity). Existing RL methods usually optimize for a binary "Pass" reward. However, the authors argue that "Pass" is relative to the test suite. Small test inputs reward brute-force; large inputs demand O(N log N) or better.

The core insight is that as you make the verifier stricter (increasing input sizes), the model doesn't just get better—it starts trading off Optimization Failures (timeouts/memory) for Correctness Failures (logic bugs).

Methodology: Mapping the Frontier

The team trained a family of checkpoints using different input-length thresholds ().

  1. Nested Rewards: rewards passing tests with input length .
  2. Linear Arithmetic: They applied weight averaging .
    • Interpolation (): Recovers the behavior between the two trained checkpoints.
    • Extrapolation ( or ): Pushes the model beyond the limits of what RL achieved.

The Correctness-Efficiency Frontier

Correctness-Efficiency Frontier Figure 1: On the left, we see that extrapolated checkpoints (alpha > 1) continue the frontier of RL training. On the right, the solve rate remains stable, but the internal composition of failures shifts.

Experiments and Results

The researchers tested this across three "Interaction Scales":

  • Pure Reasoning: Single-turn code generation.
  • Tool Use: LLM uses a Python interpreter iteratively.
  • Agentic Coding: Full sandbox access with file systems.

Across all settings and model scales (7B to 32B), the frontier held firm.

  • Extrapolation Success: For , the solve rate remained flat but the set of problems solved changed.
  • Inference-Time Ensembling: By sampling from multiple checkpoints along the extrapolated axis, they achieved a 3.3% absolute gain in pass@250 on the notoriously difficult LCB/hard benchmark.

Performance vs. Model Scale

Scale Comparison Figure 2: The frontier tracks model capability. At 32B, the trade-off is clear on "Hard" problems; at 7B, the same tension appears at the "Medium" level.

Deep Insight: Is the Model Learning New Algorithms?

The case studies (Section J) reveal a fascinating behavioral shift.

  • Low-coverage models are "lazy"—they pivot to brute-force when logic gets hard.
  • High-coverage/Extrapolated models are "ambitious"—they consistently attempt complex data structures (like Segment Trees or DP) even when they might introduce bugs.

By averaging these weights, you aren't just blending code; you are blending priors on algorithmic complexity.

Conclusion and Limitations

The study proves that the "Alignment Tax" isn't a fixed penalty but a point on a manifold. Extrapolation allows researchers to "surf" this manifold at inference time to find the sweet spot for a specific problem.

Limitations: The study is currently focused on competitive programming. Applying this to Software Engineering (SWE) tasks would require developing an "ordered axis" for test-suite strength or performance benchmarks.

Takeaway: If your RL model is plateaus, stop training. Start extrapolating.

Find Similar Papers

Try Our Examples

  • Search for recent papers using weight extrapolation or "model souping" beyond the convex hull for Reinforcement Learning from Human Feedback (RLHF) or reasoning tasks.
  • Who first proposed "Rewarded Soups" and how does the current concept of "nested verifier-induced frontiers" differ from multi-objective Pareto optimization?
  • Investigate studies that apply extrapolative weight averaging to other code-related tasks like GPU kernel optimization or Software Engineering (SWE-bench) workloads.
Contents
Extrapolative Weight Averaging: Beyond the RL Frontier in Code Generation
1. TL;DR
2. The Problem: The Hidden Trade-off in Code RL
3. Methodology: Mapping the Frontier
3.1. The Correctness-Efficiency Frontier
4. Experiments and Results
4.1. Performance vs. Model Scale
5. Deep Insight: Is the Model Learning New Algorithms?
6. Conclusion and Limitations