Agentic Boosting: How a Committee of Nano Models Can Outthink Giants

Agentic Systems as Boosting Weak Reasoning Models

2026-05-01
Varun Sunkaraneni, Pierfrancesco Beneventano, Riccardo Neumarker, Tomaso Poggio, Tomer Galanti
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a theoretical framework and empirical validation for "Agentic Boosting," demonstrating that a committee of weak models (e.g., GPT-5.4 nano) can match frontier models like Claude 4.5 by using verifier-backed search. It formalizes the orchestration of reasoning through proposal coverage and local identifiability, achieving a 76.4% solve rate on SWE-bench Verified.

TL;DR

Is the secret to better AI agents simply "more compute"? Not exactly. This paper proves that reasoning performance can be "boosted" at inference time by treating weak LLMs as a committee. By separating the act of generating a solution (coverage) from the act of recognizing it (identifiability), the researchers enabled a "nano" model to match the performance of heavyweights like Gemini 3 Pro and Claude 4.5 on the rigorous SWE-bench Verified benchmark.

Background: The Selection Gap

In the world of Large Language Models, we often evaluate success using Pass@1 (the first try) or Oracle Best-of-k (the best of many tries). There is usually a massive gap between these two metrics. This paper argues that this gap represents latent capability: the model "knows" the answer (it's in the proposal pool), but the system doesn't know how to pick it.

The authors frame this as "Inference-time Boosting," borrowing from classical machine learning theory to turn weak reasoning signals into one strong, reliable trajectory.

The Problem: The Fragility of Reasoning

Most agentic systems fail because they act like a "flat" predictor. In complex tasks like software engineering, a single mistake in a file edit can derail the entire process.

  • The Coverage Problem: Does the correct move even exist in the sampled candidates?
  • The Identifiability Problem: Even if the correct move is there, can a critic find it?

The authors prove a sobering Black-Box Separation: Simply sampling more candidates does not help you identify which one is correct. You need a separate signal—like execution tests, type checking, or specialized reviewers—to "bridge" the gap from coverage to success.

Methodology: The Committee Protocol

The paper formalizes a Committee Protocol (Πk,m,r) that operates over partial reasoning states. Instead of one model doing everything, the work is split into specialized roles:

  1. Proposers: Generate diverse candidate actions.
  2. Critics: Perform independent checks per candidate to filter out obvious bugs.
  3. Comparators: Conduct pairwise tournament votes to find the "Copeland winner" among the survivors.

Inference-time Boosting Architecture Figure 1: Comparison of the nano committee against standalone frontier models. The orchestration lifts the nano model (67%) to match Claude 4.5 level (76.4%).

The "Blind-Spot" Ceiling

A crucial insight from the paper is the Blind-Spot Floor (). If a specific type of logic is missing from a model's training data, samples won't find it. The model has a "shared blind spot." This gives us a formal way to measure the "boostable ceiling" of any given model.

Experimental Results: Matching the Giants

Using the SWE-bench Verified (500 real-world GitHub issues), the researchers put their theory to the test.

  • Baseline: GPT-5.4 nano alone solved 67.0%.
  • The Boost: Using the same nano model as proposer, critic, and comparator, the system reached 76.4%.
  • The Comparison: This "weak" committee outperformed GPT-5.4 mini and equaled the standalone performance of Gemini 3 Pro and Claude 4.5 Thinking.

Performance Scaling Figure 2: Scaling curves showing how additional proposals () and orchestration components (Critics/Comparators) close the gap toward the 79% Oracle limit.

Critical Analysis: Where do failures come from?

The paper reveals that after orchestration, most remaining failures are Coverage Failures (The correct answer wasn't in the pool) rather than Selection Failures. This means that once we have a good committee, the only way to get better is to use even more diverse proposers.

Conclusion and Future Outlook

The primary takeaway is that Information Selection is just as important as Information Generation.

For developers of agentic systems, the message is clear: Stop focusing solely on making your "Worker Agent" smarter. Instead, focus on:

  • Diversity: Use different prompts or models to eliminate "shared blind spots."
  • Local Identifiability: Build better tools (tests, linters, execution environments) that allow your agents to recognize their own success.

The future of AI agents isn't a single "God Model"—it's a perfectly orchestrated committee of specialists.

Limitations

  • Cost: Running 8 proposals and dozens of critic calls increases inference cost and latency.
  • Domain Specificity: This works best in "Verifier-backed" domains (Code, Math, Logic) where local moves can be tested. In creative writing, the "identifiability signal" is much muddier.

Find Similar Papers

Try Our Examples

  • Search for recent papers that differentiate between "proposal coverage" and "selection identifiability" in LLM-based code generation tasks.
  • Which study first introduced the concept of Process Reward Models (PRMs), and how does the current paper's "local identifiability" theory expand upon those stepwise verification methods?
  • Investigate how the "shared blind spot" (B-floor) identified in this paper affects the scaling laws of multi-agent systems in non-coding domains like medical diagnosis or legal reasoning.
Contents
Agentic Boosting: How a Committee of Nano Models Can Outthink Giants
1. TL;DR
2. Background: The Selection Gap
3. The Problem: The Fragility of Reasoning
4. Methodology: The Committee Protocol
4.1. The "Blind-Spot" Ceiling
5. Experimental Results: Matching the Giants
5.1. Critical Analysis: Where do failures come from?
6. Conclusion and Future Outlook
6.1. Limitations