GRAM: Breaking the Deterministic Ceiling of Recursive Reasoning
Generative Recursive Reasoning
This paper introduces Generative Recursive reAsoning Models (GRAM), a framework that transforms deterministic recursive reasoning into a stochastic latent-variable process. By modeling reasoning as a distribution over latent trajectories, GRAM achieves state-of-the-art results on structured reasoning tasks like Sudoku-Extreme and ARC-AGI.
TL;DR
Generative Recursive reAsoning Models (GRAM) redefine neural reasoning by shifting from fixed deterministic updates to stochastic latent trajectories. By treating reasoning as a generative process, GRAM enables models to explore multiple hypotheses in parallel, achieving superior performance on hard puzzles (Sudoku-Extreme, ARC-AGI) and multi-solution tasks where traditional models fail.
The Deterministic Trap
Current Recursive Reasoning Models (RRMs) like HRM or TRM are "deep" but "narrow." They refine a latent state through repeated applications of the same function, which is efficient for parameter count. However, their deterministic nature is a fatal flaw: once a trajectory starts toward a wrong "attractor" in the latent space, it has no mechanism to "think" its way out.
In the real world—and in complex puzzles—reasoning isn't a straight line; it's a search tree. Previous models were essentially blind to alternative branches, leading to mode collapse in tasks like N-Queens where multiple valid solutions exist.
Methodology: Stochastic Guidance & Variational Inference
GRAM solves this by making the latent transition probabilistic. Instead of a fixed , it uses: , where .
1. The Hierarchical Loop
GRAM employs a nested structure:
- Inner Loop: Rapid, deterministic refinement of low-level details.
- Outer Loop: Stochastic updates to the high-level abstract reasoning state. This is where "Stochastic Guidance" steers the trajectory.

2. Scaling: Depth vs. Width
The most significant innovation is Width-based scaling. Traditional models scale by adding more steps (Depth). GRAM can scale by sampling trajectories in parallel (Width). A "Latent Process Reward Model" (LPRM) then acts as a critic to pick the winning trajectory.
Experimental Evidence
The results prove that "Width" is often more effective than "Depth" for hard constraints.
Multi-Solution Superiority
In the N-Queens task, deterministic models (TRM/HRM) collapse, finding only one solution or failing as the number of possible solutions increases. GRAM maintains near-perfect accuracy and high coverage by sampling different paths.

Efficient Scaling
As shown in the scaling plots, a 16-iteration GRAM with 20 parallel samples (Width) crushes a 320-iteration TRM (Depth). This bypasses the sequential latency bottleneck that plagues deep recursive models.

Latent Trajectory Visualization
Visualizing the PCA-projected latent space reveals the "Why." While TRM is a single line, GRAM's 50 samples look like a search party. Some get stuck in local minima (yellow), but many find the global optimum (dark blue), a feat impossible for a deterministic agent.

Critical Analysis & Conclusion
The value of GRAM is its inductive bias: it assumes that reasoning is inherently a probabilistic search. By replacing the "greedy" deterministic update with a variational framework, it allows tiny networks (10M parameters) to punch way above their weight class, outperforming even massive LLMs on specialized logic benchmarks.
Limitations: The training relies on deep supervision and truncated backpropagation, which can be memory-intensive. Scaling this to "foundation model" size remains a significant challenge for future research.
Future Outlook: GRAM opens the door to "Generative Reasoning" where the model doesn't just predict the next token, but explores a latent forest of possibilities before committing to an answer.
