GRAM: Breaking the Deterministic Ceiling of Recursive Latent Reasoning

Generative Recursive Reasoning

2026-01-01
Junyeob Baek, Mingyu Jo, Minsu Kim, Mengye Ren, Yoshua Bengio, Sungjin Ahn
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Generative Recursive reAsoning Models (GRAM), a framework that reformulates recursive latent reasoning as a stochastic, generative process. By replacing deterministic state updates with learned stochastic transitions, GRAM achieves state-of-the-art results on structured reasoning tasks like Sudoku-Extreme and ARC-AGI.

TL;DR

Researchers have introduced Generative Recursive reAsoning Models (GRAM), a framework that shifts neural reasoning from a single, deterministic path to a probabilistic exploration of multiple latent trajectories. By treating reasoning as a generative process, GRAM allows models to scale via both depth (more steps) and width (parallel samples), achieving superior performance on ARC-AGI and Sudoku-Extreme while naturally handling problems with multiple valid solutions.

Background: The Problem with One-Way Thinking

In the quest for efficient AI reasoning, Recursive Reasoning Models (RRMs)—like Looped Transformers or TRM—have emerged as a parameter-efficient alternative to massive LLMs. Instead of generating a long "Chain-of-Thought" (CoT) sequence, they refine a persistent hidden state through repeated computation.

However, current RRMs suffer from a "deterministic collapse." Like a hiker following a single preset path, if the first few steps are slightly off, the model gets stuck in a local minimum with no way to backtrack or explore alternatives. This is why tiny recursive networks often struggle with "multi-solution" puzzles or extremely high-constraint tasks where one wrong move invalidates the entire result.

The Core Insight: Stochastic Latent Transitions

GRAM solves this by turning the latent update into a stochastic process. Instead of a fixed function , GRAM samples the next state from a distribution:

The architecture uses a clever Hierarchical Transition mechanism:

  1. Low-level (Inner Loop): A deterministic refinement that handles fine-grained local computation.
  2. High-level (Outer Loop): A stochastic update ("Stochastic Guidance") that steers the abstract reasoning trajectory.

Model Architecture

By training this system using Amortized Variational Inference, the model learns to propose diverse reasoning paths that are likely to lead to a correct solution.

Beyond Depth: Scaling the "Width"

One of the most exciting aspects of GRAM is Inference-Time Scaling. Traditionally, to make a model "smarter," you either make it bigger (parameters) or run it longer (sequential depth).

GRAM introduces a third axis: Width. Because transitions are stochastic, you can sample reasoning trajectories in parallel.

  • The Result: 20 parallel samples at 16 iterations achieve higher accuracy (97.0%) than a deterministic model running for 320 iterations (90.5%).
  • Efficiency: Parallel sampling bypasses the sequential latency bottleneck, making "thinking wider" faster than "thinking deeper" on modern GPU hardware.

Inference Scaling Comparison

Experimental Results: Slaying the Chaos

The researchers tested GRAM on several "hard" benchmarks where deterministic models typically fail:

  • Sudoku-Extreme: Puzzles with minimal clues. GRAM achieved 97% accuracy, while -mini and DeepSeek-R1 (using standard prompting) struggled to solve these specific constraint-heavy puzzles.
  • ARC-AGI: GRAM reached 52% on ARC-1, a significant jump over prior recursive SOTA.
  • Multi-Solution Tasks: In N-Queens and Graph Coloring, where multiple valid answers exist, GRAM demonstrated excellent Coverage, finding diverse solutions whereas other models would repeatedly output the same answer.

Experimental Accuracy Table

Deep Insight: Reasoning as Unconditional Generation

Remarkably, GRAM can act as an unconditional generator. If you give it an empty board, it can "reason its way" into creating a perfectly valid, unique Sudoku puzzle from scratch. This bridges the gap between structured reasoning (finding a solution) and creative generation (creating a valid structure), proving that the internal logic of a recursive model can be used to generate data that adheres to complex global constraints.

Critical Analysis & Future Outlook

While GRAM is a breakthrough for compact reasoning models, it still faces challenges:

  1. Training Efficiency: The sequential nature of deep supervision makes training slower than the massively parallel training of standard Transformers.
  2. Scalability: While 10M-parameter models perform exceptionally well, it remains to be seen if this stochastic latent approach can be scaled to the 70B+ parameter regime effectively.

The Takeaway: GRAM proves that reasoning shouldn't be a straight line. By embracing uncertainty and stochasticity in the latent space, we can build models that are not only more efficient but also more robust at solving the world's most complex logical puzzles.


Reference: Baek et al., "Generative Recursive Reasoning". Licensed under Creative Commons / MIT for respective benchmarks.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize amortized variational inference to improve the reasoning capabilities of Transformer-based architectures.
  • What are the seminal works on Recursive Reasoning Models (RRMs) and how have they evolved from Universal Transformers to current hierarchical designs?
  • Explore research that applies stochastic latent state-space models to multi-modal generative tasks beyond symbolic reasoning or puzzles.
Contents
GRAM: Breaking the Deterministic Ceiling of Recursive Latent Reasoning
1. TL;DR
2. Background: The Problem with One-Way Thinking
3. The Core Insight: Stochastic Latent Transitions
4. Beyond Depth: Scaling the "Width"
5. Experimental Results: Slaying the Chaos
6. Deep Insight: Reasoning as Unconditional Generation
7. Critical Analysis & Future Outlook