MathSmith: Forging Olympiad-Level Reasoning through Reinforced Synthetic Problems

Mathsmith: Towards extremely hard mathematical reasoning by forging synthetic problems with a reinforced policy

2025-01-01
Shaoxiong Zhan, Yanlin Lai, Ziyu Lu, Dahua Lin, Ziqing Yang, Fei Tan
Summary
Problem
Method
Results
Takeaways
Abstract

MathSmith is a novel framework for synthesizing extremely challenging mathematical problems by sampling "concept-explanation" pairs from PlanetMath and refining them through reinforcement learning. It achieves SOTA results on competition-level benchmarks, outperforming baselines by 9.8%–18.1% on AIME and OlympiadBench.

TL;DR

MathSmith is a specialized "blacksmith" for mathematical reasoning. By extracting raw concepts from PlanetMath and using reinforcement learning (RL) to forge them into complex problems, it has pushed LLMs to solve difficult AIME and Olympiad tasks. The core innovation lies in using reasoning trace length as a proxy for difficulty, ensuring that the synthesized data is not just hard, but "deep."

The Scarcity of "Hard" Math

Current Large Language Models (LLMs) are hitting a ceiling in reasoning tasks. The problem isn't just the size of the model; it's the quality of the training data. Most synthetic math datasets are simple rephrasings of the same standard problems (GSM8K/MATH). These models are often memorizing patterns rather than learning to reason from first principles.

The authors of MathSmith argue that for LLMs to reach "General Purpose Intelligence," they must move beyond human-labeled templates. They turned to PlanetMath, a repository of rigorous mathematical concepts, to provide "raw materials" that have never been seen in common training sets.

Methodology: The Three-Stage Forge

1. Raw Material Collection

Instead of starting with a question like "What is 2+2?", MathSmith starts with Concept-Explanation pairs (e.g., Quadratic Resolvent or Dynkin system). This ensures data independence and prevents the model from just "re-skinning" known problems.

2. Supervised Fine-Tuning (SFT) with Structural Constraints

The model is first taught how to "think" about problem creation. It identifies specific difficulty strategies, such as:

  • Implicit Logic: Hiding conditions that must be deduced.
  • Cross-topic Integration: Forcing the model to bridge algebra and geometry.
  • Abstract Modeling: Converting the unknown into a formalized structure.

MathSmith Workflow

3. Reinforcement Learning via GRPO

The masterpiece of MathSmith is the Reasoning-Oriented Reward. Unlike standard RL which just checks if an answer is correct, MathSmith optimizes for:

  • Reasoning Complexity: Problems that force a teacher model to produce longer CoT sequences (measured in tokens).
  • Answer Consistency: Ensuring the teacher model arrives at the same answer multiple times (determinism implies the problem is well-posed).

Experimental Battleground: Slaying the Olympiad Benchmarks

The results prove that "complexity-aware" synthesis is superior to simple "augmentation."

  • Hard Benchmarks (AIME, OlympiadBench): MathSmith variants outperformed MetaMath and NuminaMath by significant margins—up to 18.1% relative improvement on certain benchmarks.
  • Long-CoT Mastery: In the "thinking mode" (similar to OpenAI's o1 or DeepSeek-R1), MathSmith-trained models showed much higher proficiency, as the synthetic data was specifically designed to elicit long-form reasoning.

Experimental Scaling Results

Deep Insight: Why Why Long Reasoning Equals Difficulty?

The paper makes a fascinating observation: there is a direct correlation between the difficulty of a problem and the length of the reasoning trace required to solve it (as seen in Figure 2). By rewarding the generation of problems that induce long traces, MathSmith effectively automates the creation of "Hard" data without needing a human to say "this is hard."

Reasoning Trace Comparison

Critical Analysis & Takeaways

Key Contributions:

  1. Inductive Bias Toward Complexity: By using 9 strategies and RL, the model learns the essence of what makes a math problem difficult.
  2. Weakness-Focused Pipeline: A module allows for targeted generation—if a model fails at "Number Theory," MathSmith can synthesize 1,000 new Number Theory problems to patch the gap.

Limitations: The "length = difficulty" heuristic is powerful but potentially risky; a model could learn to generate "bloated" reasoning rather than "deep" reasoning (often called "length bias" in RLHF). Future iterations will likely need more granular verification of the steps within the trace.

Conclusion: MathSmith represents a shift from "Scaling Laws of Compute" to "Scaling Laws of Synthetic Intelligence." It shows that we don't just need more data; we need to forge data that forces models to think harder.

Find Similar Papers

Try Our Examples

  • Find recent papers that use Chain-of-Thought (CoT) trace length as a reward signal or proxy for cognitive complexity in reinforcement learning for LLMs.
  • How does the "Weakness-Focused" variant generation in MathSmith relate to earlier research on "Difficulty-Aware Rejection Tuning" (DART-Math) for mathematical reasoning?
  • Explore newer studies that apply Group Relative Policy Optimization (GRPO) to non-mathematical domains like code generation or law to enhance complex reasoning.
Contents
MathSmith: Forging Olympiad-Level Reasoning through Reinforced Synthetic Problems
1. TL;DR
2. The Scarcity of "Hard" Math
3. Methodology: The Three-Stage Forge
3.1. 1. Raw Material Collection
3.2. 2. Supervised Fine-Tuning (SFT) with Structural Constraints
3.3. 3. Reinforcement Learning via GRPO
4. Experimental Battleground: Slaying the Olympiad Benchmarks
5. Deep Insight: Why Why Long Reasoning Equals Difficulty?
6. Critical Analysis & Takeaways