RM-R1: Turning Reward Modeling into a Reasoning Race
Rm-r1: Reward modeling as reasoning
The paper introduces RM-R1, a family of Reasoning Reward Models (REASRMs) that transform reward modeling into a reasoning intensive task. By integrating a Chain-of-Rubrics (CoR) mechanism, RM-R1 achieves state-of-the-art performance across major benchmarks (RewardBench, RM-Bench, RMB), outperforming much larger models like GPT-4o and Llama-3.1-70B.
Executive Summary
TL;DR: Researchers from UIUC and other institutions have introduced RM-R1, a breakthrough in how we align AI with human preferences. Instead of just "feeling" which answer is better (ScalarRMs) or giving a shallow reason (GenRMs), RM-R1 thinks through the problem first—generating its own rubrics or solving the math itself—before judging. This "Reasoning Reward Model" (REASRM) approach allows a 32B model to consistently beat giants like GPT-4o and 340B-parameter models.
Background: In the LLM coordinate system, this work moves reward modeling from a simple classification task to a "thinking" task, mirroring the recent success of models like DeepSeek-R1 but applying it to the judge, not just the student.
The "Thinking" Reward Model: Why Now?
Prior alignment techniques often rely on Scalar Reward Models, which are essentially "black boxes." They provide a score but no explanation. While Generative RMs (GenRMs) added some transparency, their reasoning was often "hallucinated" post-hoc or too superficial to be useful for hard tasks in math and code.
The authors' core insight is simple: To judge a complex answer, you must first be capable of solving the problem yourself or understanding the criteria deeply. By forcing the model to categorize the prompt and generate a Chain-of-Rubrics (CoR), they ensure the judgment is grounded in logic, not just surface-level patterns.
Methodology: The Core of RM-R1
The training of RM-R1 follows a sophisticated two-stage pipeline designed to bootstrap and then refine reasoning:
- Reasoning Distillation: The model learns to "think" by imitating high-quality reasoning traces from teacher models like Claude-3.7 and OpenAI o3.
- Reinforcement Learning (GRPO): Using Group Relative Policy Optimization, the model is further tuned. Crucially, it uses verifiable rewards—the reward is based on whether the judge chose the correct ground-truth preferred response.
The Chain-of-Rubrics (CoR) Process
As shown in the architecture below, the model doesn't just output a score. It undergoes a structured rollout:
- Step 1: Categorization (Is this a Chat or Reasoning task?)
- Step 2: Deliberation (Generate rubrics for Chat; Solve the problem for Reasoning.)
- Step 3: Scoring (Evaluate candidates based on Step 2.)

Performance: Small Models, Big Impact
The results are striking. RM-R1 (32B) achieved an average score of 81.5% across major benchmarks, significantly higher than GPT-4o (77.7%) and Nemotron-4-340B (77.1%).
Key Findings from Experiments:
- Math & Code Mastery: On RM-Bench, the 32B model hit 91.8% in math, crushing the previous best of 73%.
- Scaling Law for Judges: Unlike some previous reward models where bigger wasn't always better, RM-R1 shows a clear linear trend: the more you scale the model and the inference time (more tokens for reasoning), the more accurate the judgment becomes.

Deep Insight: The Power of Warm Start
The paper provides a fascinating look at Training Dynamics. Models that started with reasoning distillation (Warm Start) were far more stable during RL than those that didn't (Cold Start). The warm-start models actually learned to be concise first before expanding their reasoning length as they became more confident—a sign of efficient internalizing of the evaluation logic.

Conclusion & Future Outlook
RM-R1 proves that Reasoning is the "Engine" of alignment. By making the reward model think, we don't just get better scores; we get a transparent, verifiable process that can scale with compute.
Takeaway: Future AI development will likely move away from simple preference datasets toward these "Self-reflecting" judges. The next frontier? Applying this reasoning-heavy judgment to multimodal tasks and autonomous agents, where the "why" behind a decision is just as important as the decision itself.
Limitations: The primary trade-off is inference latency. Reasoned judgments take more tokens and more time. However, as the authors suggest, this can be mitigated by parallelizing rollout and reward stages in the RL pipeline.
