MathSmith: Forging Olympiad-Level Reasoning via Autonomous Problem Synthesis

Mathsmith: Towards extremely hard mathematical reasoning by forging synthetic problems with a reinforced policy

2025-01-01
Shaoxiong Zhan, Yanlin Lai, Ziyu Lu, Dahua Lin, Ziqing Yang, Fei Tan
Summary
Problem
Method
Results
Takeaways
Abstract

MathSmith is a novel framework designed to synthesize extremely difficult mathematical problems by leveraging a reinforced policy and high-level concept-explanation pairs from PlanetMath. By training on these "from scratch" synthetic problems, LLMs achieved substantial performance gains, particularly in competition-level benchmarks like AIME and OlympiadBench (+9.8% to +18.1% improvement).

TL;DR

The reasoning capability of Large Language Models (LLMs) is often hit by a "data ceiling"—there simply aren't enough high-quality, ultra-hard math problems to train on. MathSmith shatters this ceiling by moving away from human-written templates. Instead, it "forges" new problems from raw mathematical concepts (sourced from PlanetMath) using a reinforced policy. By rewarding the generation of problems that force models to "think longer" (longer CoT traces), MathSmith generates data that significantly boosts performance on AIME and Olympiad-level challenges.

The Motivation: Moving Beyond Templates

Most current synthetic data methods (like MetaMath or NuminaMath) perform a "reskinning" of existing problems. While effective, this approach is inherently limited by the original human-authored distribution.

The authors of MathSmith argue for a "Bitter Lesson" approach: sustainable progress comes from compute-heavy, general-purpose methods rather than handcrafted knowledge. They propose that an agent should learn to act as a "Blacksmith"—taking raw mathematical materials and refining them into coherent, complex problems without needing a human "seed" question.

Methodology: The Three-Stage Forge

The MathSmith workflow is a sophisticated pipeline designed to ensure that synthetic problems are not just hard, but logically sound and verifiable.

  1. Concept Harvesting: Collecting 11,000 concept-explanation pairs from PlanetMath, ensuring the "raw material" is inherently advanced (e.g., Von Neumann ordinals, Dynkin systems).
  2. Supervised Fine-Tuning (SFT): Using GPT-4o to generate "cold-start" data where the model learns to combine at least two concepts into a single problem using nine predefined strategies (like Multi-step Reasoning or Cross-topic Integration).
  3. Reinforcement Learning (RL): This is the core innovation. Using Group Relative Policy Optimization (GRPO), the model is optimized based on a composite reward:
    • Complexity Reward: Driven by the token length of the reasoning trace. If a teacher model (Qwen3-30B) needs more steps to solve it, the problem is deemed more complex.
    • Consistency Reward: Ensuring that the teacher model reaches a majority consensus on the answer, which filters out nonsensical or unsolvable "hallucinated" problems.

MathSmith Workflow Figure 1: The MathSmith workflow, demonstrating the transition from raw materials to a reinforced policy.

Experimental Results: Breaking the Hard Benchmarks

The effectiveness of MathSmith is most evident in "Hard" benchmarks that represent the frontier of AI reasoning:

  • AIME 2024/2025: MathSmith-HC consistently improved accuracy by approximately 10-15% compared to baselines.
  • OlympiadBench: The gains were even more pronounced as data volume scaled. Unlike traditional methods that plateau, MathSmith's purely synthetic approach showed excellent scalability.

The Reasoning Trace Correlation

A fascinating finding is the correlation between problem difficulty and reasoning length. As shown in the chart below, problems forged by MathSmith (HC and Hard versions) elicit significantly longer reasoning traces from teacher models compared to established datasets like GSM8K or even MATH.

Reasoning Trace Comparison Figure 2: Average reasoning trace length across datasets. MathSmith variants induce the deepest "thinking" mode.

Targeted Improvement: The Weakness-Focused Pipeline

One of the most practical features of MathSmith is its ability to perform "surgery" on model weaknesses. Because every problem is tied to specific source concepts, researchers can identify where a model fails (e.g., Lattices with Operators) and command MathSmith to forge a "Practice Set" of 1,000 variants targeting that specific gap. This iterative refinement led to consistent accuracy jumps on identified weak spots.

Critical Analysis & Future Outlook

While MathSmith is a breakthrough in synthetic data generation, it faces a few challenges:

  • Heuristic Limitations: Using "length of reasoning" as a proxy for difficulty is clever but potentially susceptible to "reward hacking" where a model generates unnecessarily verbose but logically simple problems.
  • Domain Specificity: Currently focused on math, the "concept-to-problem" forging logic needs significant adaptation to work in more subjective domains like creative writing or law.

Conclusion: MathSmith proves that we don't need more human problems to make models smarter; we need better "Blacksmiths" to forge the training data of the future. By focusing on reasoning depth as a trainable objective, this work paves the way for LLMs to tackle increasingly abstract and non-standard mathematical frontiers.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize "reasoning trace length" or "inference-time compute scaling" as a reward signal for Reinforcement Learning from Human Feedback (RLHF) or RLAIF.
  • What are the primary theoretical limitations of using synthetic data from PlanetMath or similar encyclopedias for mathematical problem generation compared to human-curated datasets?
  • How can the "Weakness-Focused Improvement Pipeline" from MathSmith be adapted for other specialized domains like code generation or law where concept-driven synthesis is required?
Contents
MathSmith: Forging Olympiad-Level Reasoning via Autonomous Problem Synthesis
1. TL;DR
2. The Motivation: Moving Beyond Templates
3. Methodology: The Three-Stage Forge
4. Experimental Results: Breaking the Hard Benchmarks
4.1. The Reasoning Trace Correlation
5. Targeted Improvement: The Weakness-Focused Pipeline
6. Critical Analysis & Future Outlook