GRAM: Breaking the Deterministic Ceiling of Recursive Reasoning

Generative Recursive Reasoning

2026-01-01
Junyeob Baek, Mingyu Jo, Minsu Kim, Mengye Ren, Yoshua Bengio, Sungjin Ahn
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Generative Recursive reAsoning Models (GRAM), a framework that transforms deterministic recursive reasoning into a stochastic latent-variable process. By modeling reasoning as a distribution over latent trajectories, GRAM achieves state-of-the-art results on structured reasoning tasks like Sudoku-Extreme and ARC-AGI.

TL;DR

Generative Recursive reAsoning Models (GRAM) redefine neural reasoning by shifting from fixed deterministic updates to stochastic latent trajectories. By treating reasoning as a generative process, GRAM enables models to explore multiple hypotheses in parallel, achieving superior performance on hard puzzles (Sudoku-Extreme, ARC-AGI) and multi-solution tasks where traditional models fail.

The Deterministic Trap

Current Recursive Reasoning Models (RRMs) like HRM or TRM are "deep" but "narrow." They refine a latent state through repeated applications of the same function, which is efficient for parameter count. However, their deterministic nature is a fatal flaw: once a trajectory starts toward a wrong "attractor" in the latent space, it has no mechanism to "think" its way out.

In the real world—and in complex puzzles—reasoning isn't a straight line; it's a search tree. Previous models were essentially blind to alternative branches, leading to mode collapse in tasks like N-Queens where multiple valid solutions exist.

Methodology: Stochastic Guidance & Variational Inference

GRAM solves this by making the latent transition probabilistic. Instead of a fixed , it uses: , where .

1. The Hierarchical Loop

GRAM employs a nested structure:

  • Inner Loop: Rapid, deterministic refinement of low-level details.
  • Outer Loop: Stochastic updates to the high-level abstract reasoning state. This is where "Stochastic Guidance" steers the trajectory.

GRAM Architecture

2. Scaling: Depth vs. Width

The most significant innovation is Width-based scaling. Traditional models scale by adding more steps (Depth). GRAM can scale by sampling trajectories in parallel (Width). A "Latent Process Reward Model" (LPRM) then acts as a critic to pick the winning trajectory.

Experimental Evidence

The results prove that "Width" is often more effective than "Depth" for hard constraints.

Multi-Solution Superiority

In the N-Queens task, deterministic models (TRM/HRM) collapse, finding only one solution or failing as the number of possible solutions increases. GRAM maintains near-perfect accuracy and high coverage by sampling different paths.

Performance Comparison

Efficient Scaling

As shown in the scaling plots, a 16-iteration GRAM with 20 parallel samples (Width) crushes a 320-iteration TRM (Depth). This bypasses the sequential latency bottleneck that plagues deep recursive models.

Inference Scaling

Latent Trajectory Visualization

Visualizing the PCA-projected latent space reveals the "Why." While TRM is a single line, GRAM's 50 samples look like a search party. Some get stuck in local minima (yellow), but many find the global optimum (dark blue), a feat impossible for a deterministic agent.

Latent Trajectories

Critical Analysis & Conclusion

The value of GRAM is its inductive bias: it assumes that reasoning is inherently a probabilistic search. By replacing the "greedy" deterministic update with a variational framework, it allows tiny networks (10M parameters) to punch way above their weight class, outperforming even massive LLMs on specialized logic benchmarks.

Limitations: The training relies on deep supervision and truncated backpropagation, which can be memory-intensive. Scaling this to "foundation model" size remains a significant challenge for future research.

Future Outlook: GRAM opens the door to "Generative Reasoning" where the model doesn't just predict the next token, but explores a latent forest of possibilities before committing to an answer.

Find Similar Papers

Try Our Examples

  • Search for recent papers attempting to solve the ARC-AGI challenge using latent reasoning or test-time compute scaling instead of large-scale pretraining.
  • Which paper first proposed the concept of "Looped Transformers," and how does the introduction of stochastic latent states in GRAM specifically address the convergence issues noted in that work?
  • Explore whether the GRAM framework of stochastic latent transitions has been applied to generative protein design or chemical synthesis tasks where multi-step constraint satisfaction is essential.
Contents
GRAM: Breaking the Deterministic Ceiling of Recursive Reasoning
1. TL;DR
2. The Deterministic Trap
3. Methodology: Stochastic Guidance & Variational Inference
3.1. 1. The Hierarchical Loop
3.2. 2. Scaling: Depth vs. Width
4. Experimental Evidence
4.1. Multi-Solution Superiority
4.2. Efficient Scaling
5. Latent Trajectory Visualization
6. Critical Analysis & Conclusion