RM-R1: Turning Reward Modeling into a Reasoning Powerhouse

Rm-r1: Reward modeling as reasoning

2025-01-01
Xiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin, Cheng Qian, Yu Wang, Hongru Wang, Yu Zhang, Denghui Zhang, Tong Zhang, Hanghang Tong, Heng Ji
Summary
Problem
Method
Results
Takeaways
Abstract

RM-R1 is a novel class of Reasoning Reward Models (ReasRMs) that reformulates reward modeling as a reasoning task. By integrating a Chain-of-Rubrics (CoR) mechanism and a two-stage training pipeline (distillation and RLVR), the RM-R1 family (7B to 32B) achieves state-of-the-art performance, outperforming much larger models like Llama-3.1-70B and proprietary ones like GPT-4o on major benchmarks.

Executive Summary

RM-R1 introduces a paradigm shift in how we align Large Language Models: Reward Modeling as Reasoning. While traditional reward models (RMs) act as "black-box" scorers, RM-R1 treats evaluation as a deliberate cognitive task. By using a Chain-of-Rubrics (CoR) mechanism and a specialized training recipe involving reasoning distillation and Reinforcement Learning (RL), RM-R1 achieves state-of-the-art results. The 32B version of RM-R1 consistently outperforms massive 70B and 340B models, proving that how a model thinks about a preference is more important than how many parameters it has.

The Problem: The Transparency Gap in Reward Modeling

Existing RMs fall into two camps:

  1. ScalarRMs: Efficient but opaque. They output a single number, providing zero insight into why one response was preferred.
  2. GenRMs: Transparent but often "dumb." They can explain their logic, but the reasoning is typically shallow and fails to catch subtle errors in complex math or coding tasks.

The authors argue that accurate reward signals require Deep Thinking. To judge a pair of responses, a model must infer the user's intent, resolve conflicting criteria (e.g., helpfulness vs. safety), and—in technical domains—verify the factual correctness of the solution.

Methodology: The Chain-of-Rubrics (CoR)

The core innovation of RM-R1 is the Chain-of-Rubrics. Instead of a one-size-fits-all evaluation, RM-R1 adapts its internal "thought process" based on the task type.

1. Task Perception & Strategy

  • Chat Tasks: The model generates a set of customized rubrics (e.g., politeness, nuance, emotional safety) and justifies them before evaluating.
  • Reasoning Tasks: The model first solves the problem itself (generating an internal solution) and then uses that solution as a reference to grade the candidates.

2. Two-Stage Training Pipeline

  • Stage 1: Reasoning Distillation: The model is fine-tuned on ~9k high-quality reasoning traces "bootstrapped" from frontier models like OpenAI o3 and Claude-3.7-Sonnet.
  • Stage 2: RL with Verifiable Rewards (RLVR): Using Group Relative Policy Optimization (GRPO), the model is trained to maximize the probability of picking the correct "chosen" response, effectively learning to "search" for the best logical path to the right judgment.

RM-R1 Training Pipeline

Experimental Battleground: Small Models, Big Impact

RM-R1 was tested on three brutal benchmarks: RewardBench, RM-Bench, and RMB.

CategoryModelAverage Performance
ScalarRMINF-ORM-Llama3.1-70B78.8%
GenRMGPT-4o77.7%
ReasRMRM-R1-32B81.5%

Key Insights from Results:

  • The Scaling Law of Reasoning: Unlike standard RMs where 7B models sometimes beat 27B models, RM-R1 shows a linear scaling benefit. Larger models are better at leveraging reasoning traces to improve their judgments.
  • Inference-Time Compute: RM-R1's performance increases as you allow it more "thinking tokens" during inference. More tokens = deeper reasoning = better rewards.
  • Reasoning vs. SFT: Training on reasoning chains (Distilled + RL) provided a massive +4% gain over simply training the model on answer labels (SFT).

Performance Scaling

Deep Insight: Why Does RL Help?

A fascinating finding in the paper is the difference between "Cold Start RL" and "Warm Start RL."

  • Cold Start: The model starts with short responses and slowly "learns" to reason, but often becomes unstable or overfits.
  • Warm Start (RM-R1): By distilling reasoning before RL, the model starts with a strong baseline. RL then acts as a "critical thinker," refining the model to prioritize high-impact rubrics (like medical accuracy) over superficial features (like response length).

Conclusion & Future Outlook

RM-R1 proves that the "Reasoning Model" (R1-style) paradigm isn't just for solving math problems—it's the future of model evaluation.

The implications are clear: To build the next generation of SOTA LLMs, we need Reward Models that can think, justify, and solve the very problems they are grading. Future work will likely extend this to multimodal tasks and "active preference collection," where the RM tells humans when it needs more guidance to build a better rubric.

Takeaways for Researchers:

  1. Don't skip distillation: Seeding a model with 9k high-quality reasoning traces is more effective than RL alone.
  2. Task Categorization matters: A math problem requires a different judging mindset than a safety query.
  3. Transparency is performance: More interpretable rubrics lead to more accurate final scores.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Reinforcement Learning with Verifiable Rewards (RLVR) for non-mathematical tasks like safety or creative writing.
  • Which paper originally introduced the concept of generative reward models, and how does RM-R1's "Chain-of-Rubrics" specifically differ from previous critique-based approaches?
  • Investigate studies that explore the "Inference-time Scaling Law" specifically for reward models or evaluators in LLM post-training.
Contents
RM-R1: Turning Reward Modeling into a Reasoning Powerhouse
1. Executive Summary
2. The Problem: The Transparency Gap in Reward Modeling
3. Methodology: The Chain-of-Rubrics (CoR)
3.1. 1. Task Perception & Strategy
3.2. 2. Two-Stage Training Pipeline
4. Experimental Battleground: Small Models, Big Impact
4.1. Key Insights from Results:
5. Deep Insight: Why Does RL Help?
6. Conclusion & Future Outlook
6.1. Takeaways for Researchers: