RM-R1: Redefining Reward Modeling as a Reasoning Process
Rm-r1: Reward modeling as reasoning
RM-R1 is a novel family of Reasoning Reward Models (ReasRMs) that treat reward modeling as a reasoning task using a Chain-of-Rubrics (CoR) mechanism. By integrating long-form reasoning traces before preference judgment, RM-R1 models (7B-32B) achieve state-of-the-art performance on RewardBench, RM-Bench, and RMB, outperforming much larger models like Llama-3.1-70B and proprietary models like GPT-4o.
TL;DR
Reward modeling is the backbone of aligning AI with human preferences, but current methods are often opaque or superficial. RM-R1 transforms reward modeling into a "thinking" task. By self-generating evaluation rubrics and solving problems before judging them, RM-R1 reaches new SOTA heights—beating GPT-4o and 70B models with only a 32B parameter footprint.
Background: The Hidden Weakness of Reward Models
In the typical RLHF (Reinforcement Learning from Human Feedback) pipeline, a Reward Model (RM) acts as a proxy for human judgment. However, existing RMs suffer from two major flaws:
- Scalar RMs are Black Boxes: They output a single score but don't explain why, making them hard to debug and trust.
- Generative RMs are Shallow: While they can write "critiques," they often hallucinate or focus on surface features like length and politeness rather than actual logic.
RM-R1 argues that to judge a response accurately, the model must reason about it first.
Methodology: The Chain-of-Rubrics (CoR)
The core innovation of RM-R1 is the Chain-of-Rubrics (CoR). Instead of jumping straight to a score, the model follows a structured cognitive path:
- Task Categorization: The model identifies if the prompt is a general "Chat" task or a "Reasoning" (Math/Code) task.
- Strategy Adaptation:
- For Chat: It generates specific rubrics (e.g., "Is the medical advice factually accurate?") and justifies them.
- For Reasoning: It solves the problem itself first to establish a "ground truth solution" before looking at the candidate's answer.
- Final Judgment: It compares candidates against the self-generated rubrics or solution.
The Training Pipeline
The authors used a two-step process to bake this capability into the model:
- Step 1: Distillation: Training on ~9k high-quality reasoning traces generated by "teacher" models like OpenAI's o3.
- Step 2: RL Training (GRPO): Using reinforcement learning to maximize the accuracy of the final judgment based on verifiable preference labels.

SOTA Results: Smaller Models, Better Judgments
RM-R1 doesn't just provide better explanations; it is fundamentally more accurate. On RM-Bench (the most reasoning-intensive benchmark), the 32B version of RM-R1 crushed the competition.
| Model Category | Model Name | RewardBench | RM-Bench | RMB |
|---|---|---|---|---|
| Proprietary | GPT-4o | 86.7 | 72.5 | 73.8 |
| Large Open Scalar | INF-ORM-Llama3.1-70B | 95.1 | 70.9 | 70.5 |
| Our Method | RM-R1-32B | 91.4 | 83.9 | 73.0 |
The takeaway is clear: Reasoning is a force multiplier for accuracy. RM-R1-14B already outperforms Llama-3.1-70B-Instruct despite being 5x smaller.
The Scaling Law of Reasoning
One of the most exciting findings in the paper is that Reasoning Reward Models scale better. As you increase the number of reasoning tokens at inference time, the accuracy continues to rise. This confirms that "thinking longer" makes for a better judge.

Critical Insight: Why Does This Work?
The authors provide a formal proof (Proposition 1) showing that SFT (Supervised Fine-Tuning) often fails because models learn "trivial shortcuts" (like always picking the longer answer). Reinforcement Learning (RL) is necessary because it forces the model to explore "disagreement events" where the trivial shortcut and the robust logic path provide different results. By exploring these during training, RM-R1 learns to ignore the shortcut and follow the logic.
Conclusion and Future Impact
RM-R1 represents a shift toward interpretable alignment. In the future, we shouldn't just ask if an AI is "helpful"; we should ask it to show its work. This technology paves the way for "active preference collection," where reward models can tell humans exactly what rubrics they are using and ask for clarification, making AI safer and more aligned with complex human values.
Limitations
- Latency: Thinking takes time. Reasoning RMs are slower than scalar ones.
- Verification: While reasoning tasks are easy to verify, "Chat" rubrics still rely on the model's internal understanding of "politeness" or "safety."
Senior Academic Tech Editor's Note: RM-R1 proves that the "Next-Token Prediction" paradigm is evolving into a "Next-Thought Prediction" paradigm, even for the components that evaluate the AI itself.
