GRAM: Turning Recursive Reasoning into a Generative Exploration of Latent Space

Generative Recursive Reasoning

2026-01-01
Junyeob Baek, Mingyu Jo, Minsu Kim, Mengye Ren, Yoshua Bengio, Sungjin Ahn
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Generative Recursive reAsoning Models (GRAM), a framework that transforms deterministic recursive reasoning into a probabilistic generative process. By modeling reasoning as a stochastic latent trajectory, GRAM achieves state-of-the-art results on structured tasks like Sudoku-Extreme and ARC-AGI, outperforming prior recursive models through parallel trajectory sampling and multi-hypothesis exploration.

TL;DR

The future of AI reasoning might not just be about making models "deeper" (more layers or steps), but also making them "wider" (exploring more paths). GRAM (Generative Recursive reAsoning Models) moves away from the deterministic "one-path-only" approach of current recursive models. By treating reasoning as a stochastic latent trajectory, GRAM can explore multiple hypotheses in parallel, successfully solving ultra-hard puzzles like Sudoku-Extreme and ARC-AGI while outperforming much larger models.

The Problem: The Deterministic Trap of Recursive Reasoning

Recursive Reasoning Models (RRMs) are fascinating because they decouple reasoning depth from parameter count. Instead of a 100-layer model, you use a 1-layer model 100 times. However, existing RRMs (like TRM or HRM) have a fatal flaw: they are deterministic.

Given a specific input, a deterministic model will always follow the exact same refinement path. In complex reasoning, this is equivalent to a human being unable to "change their mind" or "try a different angle" if they get stuck. This leads to mode collapse, where the model cannot find alternative solutions if the first path it takes is a dead end.

Methodology: Stochastic Guidance and Latent Exploration

GRAM solves this by reformulating the reasoning process as a probabilistic generative model.

1. Stochastic Latent Transitions

Rather than a simple , GRAM introduces Stochastic Guidance (). At each step, the model generates a deterministic "proposal" , and then adds a learned Gaussian perturbation: This allows the model to "jump" to different regions of the latent space, exploring various reasoning branches simultaneously.

2. Hierarchical Architecture

GRAM utilizes a two-level hierarchy. A low-level component handles fine-grained, deterministic intermediate computation, while a high-level component manages abstract reasoning and introduces stochasticity.

Model Architecture Figure: The GRAM architecture showing the hierarchy between low-level refinement and high-level stochastic updates.

3. Width-Based Scaling

Because GRAM is a generative model, it can scale during inference not just by running for more steps (depth), but by sampling multiple trajectories in parallel (width). To pick the winner, it uses a Latent Process Reward Model (LPRM)—a specialized head that predicts which trajectory is most likely to produce a correct answer.

Experimental Showdown: Puzzles & Multi-Solution Tasks

The researchers tested GRAM on benchmarks designed to break standard models:

  • Sudoku-Extreme: Puzzles with minimal clues.
  • ARC-AGI: Visual abstract reasoning.
  • N-Queens/Graph Coloring: Tasks with multiple valid answers.

Key Result: Parallel Scaling vs. Depth Scaling

One of the most striking findings is that width (parallel sampling) is often more efficient than depth (sequential steps). A GRAM model with 20 parallel samples at 16 steps outperformed the deterministic TRM baseline even when the latter was given 320 steps.

Performance Comparison Figure: GRAM scales both with iterations (Depth) and samples (Width), allowing it to bypass the traditional latency bottlenecks of deep recursion.

Multi-Solution Coverage

In tasks like N-Queens where many solutions are possible, deterministic models like TRM and HRM see their accuracy plummet as the number of possible solutions increases. GRAM, however, maintains flat, consistent performance because it can "see" and explore the whole solution landscape.

Unconditional Generation: Reasoning as Creation

Perhaps surprisingly, GRAM can also function as a standard generative model. When given a blank board, it can generate valid, unique Sudoku puzzles with 99.05% validity. This is far superior to specialized discrete diffusion models (D3PM), primarily because GRAM’s recursive refinement allows it to correct its own "hallucinations" as it generates.

Sudoku Generation Figure: Qualitative examples of Sudoku boards generated "from thin air" by GRAM, demonstrating perfect constraint satisfaction.

Deep Insight: Navigating the Latent Landscape

Visualization of the latent space via PCA shows a clear difference: TRM is a single line that might get stuck in a "yellow" (high error) zone. GRAM represents a "cloud" of 50 trajectories. Even if some get stuck, others find the "dark blue" (global optimum) path.

Conclusion & Future Outlook

GRAM demonstrates that for a model to be an expert reasoner, it must be able to explore. By marrying Recursive Architectures with Variational Inference, the authors have created a framework where reasoning is no longer a straight line, but a dynamic search process.

Limitations: The primary bottleneck remains training efficiency; recursive deep supervision is computationally expensive compared to the massive parallelism of standard Transformers. However, as we look for ways to make "smaller" models smarter at the "edge," GRAM’s architecture offers a compelling path forward.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply variational inference or stochastic latent variables to improve the "Chain of Thought" reasoning efficiency in Transformers.
  • Find the original research on "Adaptive Computation Time" (ACT) and investigate how GRAM's stochastic guidance differs from earlier probabilistic RNN approaches like VRNN.
  • Explore newer studies that use Latent Process Reward Models or value-based trajectory selection to scale test-time compute in LLM reasoning tasks.
Contents
GRAM: Turning Recursive Reasoning into a Generative Exploration of Latent Space
1. TL;DR
2. The Problem: The Deterministic Trap of Recursive Reasoning
3. Methodology: Stochastic Guidance and Latent Exploration
3.1. 1. Stochastic Latent Transitions
3.2. 2. Hierarchical Architecture
3.3. 3. Width-Based Scaling
4. Experimental Showdown: Puzzles & Multi-Solution Tasks
4.1. Key Result: Parallel Scaling vs. Depth Scaling
4.2. Multi-Solution Coverage
5. Unconditional Generation: Reasoning as Creation
6. Deep Insight: Navigating the Latent Landscape
7. Conclusion & Future Outlook