STV: Breaking the Self-Improvement Ceiling with Self-Trained Verification
Self-Trained Verification for Training- and Test-Time Self-Improvement
This paper introduces Self-Trained Verification (STV), a method to scale reasoning performance by training verifiers to catch their own errors using reference-conditioned distillation. It achieves SOTA gains in math and science tasks, such as lifting SciKnowEval performance from 1.5% to 21.0% using Qwen3-8B.
TL;DR
Reasoning models often plateau because they "don't know what they don't know." Self-Trained Verification (STV) breaks this by training verifiers to catch self-generated errors using a simple trick: a model is much better at diagnosing a mistake if it can look at the answer key. By distilling this "answer-key-aware" diagnostic skill back into the base model, the authors achieve massive gains—up to 14x on scientific reasoning—and show that better verification actually makes for better standalone generators.
The Verification Bottleneck
In the current LLM landscape, "Self-Improvement" typically happens in two places:
- Test-time: Verification-Refinement (V-R) loops where the model checks its work and tries again.
- Training-time: Self-training where the model learns from its own successful attempts.
Both are capped by the Verifier. Current verifiers are notoriously unreliable; they suffer from "reward hacking" (giving high scores to plausible but wrong answers) and provide hollow feedback like "Your answer might be wrong." Without a precise signal to identify where a logic gate failed, more compute just leads to more confident hallucinations.
Methodology: The Asymmetry of Diagnosis
The core insight of STV is an informational asymmetry: diagnosing a flaw is hard, but comparing a flawed solution to a correct reference is significantly easier.
The authors use a three-step training process:
- Teacher Generation: A "Teacher" verifier is prompted with the problem, the student's attempt, and the reference solution. This teacher can easily pinpoint the exact logical gap.
- On-Policy Distillation (OPD): A "Student" verifier (without the reference) is trained to mimic the teacher's diagnostic distribution. This forces the student to learn the features of a mistake.
- Verdict RL: The verifier is further tuned with Reinforcement Learning to ensure its binary "Accept/Reject" labels match the ground truth.

Beyond Scaling: Verifier-in-the-Loop (ViL) Training
The paper doesn't stop at test-time fixes. They introduce Verifier-in-the-Loop (ViL) training. Instead of training the generator on static datasets, they train it inside the V-R loop. The generator learns specifically how to act on the STV verifier's feedback.
The surprising result? Training-time self-improvement. Even when the verifier is removed at test time (Round 0), the ViL-trained generator performs 30% better than a model trained with standard RL (RLVR). This suggests that learning to process diagnostic feedback actually instills deeper "first-try" reasoning capabilities.
Experimental Breakthroughs
The results on hard reasoning benchmarks demonstrate that trained verification can substitute for model scale:
- Math (DAPO): On the "Hardest" split, 8B models guided by STV doubled their accuracy, outperforming the 32B model.
- Science (SciKnowEval): A massive jump from 1.5% to 21.0% on the hardest problems.
- Calibration: Unlike untrained verifiers, STV verifiers show a strong correlation between their scores and actual accuracy, effectively mitigating reward hacking.

Critical Analysis & Conclusion
Takeaway
STV proves that "learning to verify" is just as important as "learning to solve." By turning the reference solution into a supervision signal for feedback quality, the authors have provided a scalable path for models to climb the reasoning ladder without human-annotated critiques.
Limitations
- Reference Dependence: The training requires ground-truth solutions. In truly novel frontier research (e.g., unsolved math), this signal isn't available, necessitating future work on "unsupervised" verification.
- Compute Costs: While efficient compared to scaling model size, running 20+ rounds of refinement is still significantly more expensive than a single pass.
Future Outlook
The next frontier in AI reasoning will likely involve iterative cycles where a model spends half its training budget learning to be a critic. As verifiers get stronger, they allow the generator to explore more complex problem spaces, creating a virtuous cycle of self-improvement that could eventually move beyond human-level reasoning.
