GRAM: Breaking the Deterministic Barrier in Recursive Reasoning

Generative Recursive Reasoning

2026-01-01
Junyeob Baek, Mingyu Jo, Minsu Kim, Mengye Ren, Yoshua Bengio, Sungjin Ahn
Summary
Problem
Method
Results
Takeaways
Abstract

Generative Recursive reAsoning Models (GRAM) is a novel framework that transforms deterministic recursive reasoning into a probabilistic generative process. By introducing stochastic latent transitions, GRAM achieves SOTA performance on structured reasoning benchmarks, notably reaching 97.0% on Sudoku-Extreme and outperforming prior recursive baselines like TRM and HRM.

TL;DR

Generative Recursive reAsoning Models (GRAM) represents a paradigm shift from deterministic to probabilistic latent reasoning. While previous recursive models followed a single, fixed path to a solution, GRAM utilizes stochastic latent transitions to explore multiple reasoning trajectories simultaneously. This allows for superior performance on complex puzzles (Sudoku-Extreme), multi-solution discovery (N-Queens), and even unconditional generation from empty inputs.

Problem & Motivation: The Deterministic Dead-End

Recursive Reasoning Models (RRMs) like HRM and TRM have shown that we can increase a model's "intelligence" by simply applying the same weights over and over—essentially thinking longer without adding parameters. However, these models were deterministic.

In complex reasoning, you often hit a "fork in the road." A deterministic model is forced to pick one path and stay on it. If that path is a local minimum, the model fails. Prior work lacked the "imagination" to maintain multiple hypotheses or backtrack through latent exploration. GRAM's core insight is that reasoning should be modeled as a stochastic process, not a straight line.

Methodology: High-Level Stochastic Guidance

GRAM's architecture is a hierarchical recycler. It splits the latent state into a high-level (h) and low-level (l) component.

  1. Inner Loop: The low-level state is refined deterministically to handle fine-grained computation.
  2. Outer Loop: The high-level state receives a Stochastic Guidance (SG) signal. Instead of a fixed update, the model samples:

This essentially "nudges" the reasoning trajectory, allowing the model to jump out of local minima.

GRAM Architecture Figure: The GRAM architecture showcasing the interaction between deterministic proposals () and stochastic guidance ().

Experiments & Results: Width vs. Depth

One of the most profound findings of GRAM is Inference-Time Scaling. Traditionally, more compute meant more depth (iterations). GRAM introduces width (parallel samples).

1. Scaling Success

The researchers found that sampling multiple trajectories in parallel (and picking the best via a Reward Model) is more efficient than just going deeper. GRAM with 20 parallel samples at 16 steps crushed TRM running at 320 steps.

Inference-Time Scaling Figure: Performance improvements on Sudoku-Extreme scaling with both depth (iterations) and width (N samples).

2. Multi-Solution Coverage

In N-Queens or Graph Coloring, there isn't just one "right" answer. Deterministic models collapse (mode collapse), capturing only ~36% of solutions. GRAM, through its stochastic nature, explores the solution space and achieves nearly 90% coverage.

3. Unconditional Generation

Remarkably, GRAM can generate valid Sudoku boards and MNIST digits starting from a blank grid. This proves it isn't just a solver—it's a true generative model that understands the underlying distribution of constraints.

Latent Trajectories Figure: Visualization of 50 sampled trajectories in the latent space. While TRM (deterministic) follows one path, GRAM explores the landscape to find the global optimum.

Critical Analysis & Conclusion

Takeaway

GRAM proves that probabilistic recursion is a fundamental design principle for future reasoning architectures. It decouples the "thinking" process from the token-generation process, allowing for compact, efficient models that can explore complex solution spaces.

Limitations

The primary bottleneck is the sequential nature of deep supervision. Training these models is memory-intensive and slow because you must backpropagate through many steps of the recursion, making it harder to scale to "Foundation Model" sizes compared to standard parallel Transformers.

Future Work

The next frontier is applying GRAM's principles to Multi-modal Large Language Models. Imagine an LLM that doesn't just predict the next token, but stochastically refines a "latent thought" until it reaches high confidence, scaling its "thinking time" dynamically based on task difficulty.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend the concept of "latent reasoning" or "internal chain-of-thought" beyond the Transformer architecture, specifically those incorporating stochasticity.
  • Which original studies established the "Looped Transformer" or "Universal Transformer" as the basis for parameter-efficient depth, and how does GRAM's variational approach fundamentally alter their mathematical objective?
  • Search for research applying state-space models (SSMs) or diffusion-based refinement to combinatorial optimization and constraint satisfaction tasks similar to Sudoku and Graph Coloring.
Contents
GRAM: Breaking the Deterministic Barrier in Recursive Reasoning
1. TL;DR
2. Problem & Motivation: The Deterministic Dead-End
3. Methodology: High-Level Stochastic Guidance
4. Experiments & Results: Width vs. Depth
4.1. 1. Scaling Success
4.2. 2. Multi-Solution Coverage
4.3. 3. Unconditional Generation
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Work