MathSmith: Forging Olympiad-Level Reasoning through Reinforced Synthetic Problems
Mathsmith: Towards extremely hard mathematical reasoning by forging synthetic problems with a reinforced policy
MathSmith is a novel framework for synthesizing extremely challenging mathematical problems by sampling "concept-explanation" pairs from PlanetMath and refining them through reinforcement learning. It achieves SOTA results on competition-level benchmarks, outperforming baselines by 9.8%–18.1% on AIME and OlympiadBench.
TL;DR
MathSmith is a specialized "blacksmith" for mathematical reasoning. By extracting raw concepts from PlanetMath and using reinforcement learning (RL) to forge them into complex problems, it has pushed LLMs to solve difficult AIME and Olympiad tasks. The core innovation lies in using reasoning trace length as a proxy for difficulty, ensuring that the synthesized data is not just hard, but "deep."
The Scarcity of "Hard" Math
Current Large Language Models (LLMs) are hitting a ceiling in reasoning tasks. The problem isn't just the size of the model; it's the quality of the training data. Most synthetic math datasets are simple rephrasings of the same standard problems (GSM8K/MATH). These models are often memorizing patterns rather than learning to reason from first principles.
The authors of MathSmith argue that for LLMs to reach "General Purpose Intelligence," they must move beyond human-labeled templates. They turned to PlanetMath, a repository of rigorous mathematical concepts, to provide "raw materials" that have never been seen in common training sets.
Methodology: The Three-Stage Forge
1. Raw Material Collection
Instead of starting with a question like "What is 2+2?", MathSmith starts with Concept-Explanation pairs (e.g., Quadratic Resolvent or Dynkin system). This ensures data independence and prevents the model from just "re-skinning" known problems.
2. Supervised Fine-Tuning (SFT) with Structural Constraints
The model is first taught how to "think" about problem creation. It identifies specific difficulty strategies, such as:
- Implicit Logic: Hiding conditions that must be deduced.
- Cross-topic Integration: Forcing the model to bridge algebra and geometry.
- Abstract Modeling: Converting the unknown into a formalized structure.

3. Reinforcement Learning via GRPO
The masterpiece of MathSmith is the Reasoning-Oriented Reward. Unlike standard RL which just checks if an answer is correct, MathSmith optimizes for:
- Reasoning Complexity: Problems that force a teacher model to produce longer CoT sequences (measured in tokens).
- Answer Consistency: Ensuring the teacher model arrives at the same answer multiple times (determinism implies the problem is well-posed).
Experimental Battleground: Slaying the Olympiad Benchmarks
The results prove that "complexity-aware" synthesis is superior to simple "augmentation."
- Hard Benchmarks (AIME, OlympiadBench): MathSmith variants outperformed MetaMath and NuminaMath by significant margins—up to 18.1% relative improvement on certain benchmarks.
- Long-CoT Mastery: In the "thinking mode" (similar to OpenAI's o1 or DeepSeek-R1), MathSmith-trained models showed much higher proficiency, as the synthetic data was specifically designed to elicit long-form reasoning.

Deep Insight: Why Why Long Reasoning Equals Difficulty?
The paper makes a fascinating observation: there is a direct correlation between the difficulty of a problem and the length of the reasoning trace required to solve it (as seen in Figure 2). By rewarding the generation of problems that induce long traces, MathSmith effectively automates the creation of "Hard" data without needing a human to say "this is hard."

Critical Analysis & Takeaways
Key Contributions:
- Inductive Bias Toward Complexity: By using 9 strategies and RL, the model learns the essence of what makes a math problem difficult.
- Weakness-Focused Pipeline: A module allows for targeted generation—if a model fails at "Number Theory," MathSmith can synthesize 1,000 new Number Theory problems to patch the gap.
Limitations: The "length = difficulty" heuristic is powerful but potentially risky; a model could learn to generate "bloated" reasoning rather than "deep" reasoning (often called "length bias" in RLHF). Future iterations will likely need more granular verification of the steps within the trace.
Conclusion: MathSmith represents a shift from "Scaling Laws of Compute" to "Scaling Laws of Synthetic Intelligence." It shows that we don't just need more data; we need to forge data that forces models to think harder.
