GRAM: Transforming Recursive Reasoning into Probabilistic Latent Exploration
Generative Recursive Reasoning
This paper introduces Generative Recursive reAsoning Models (GRAM), a framework that transforms deterministic recursive reasoning into a stochastic latent-variable process. By modeling reasoning as a distribution over trajectories, it achieves SOTA results on Sudoku-Extreme (97.0%) and ARC-AGI while enabling multi-hypothesis exploration.
TL;DR
Researchers have moved beyond deterministic "refinement" in neural reasoning. GRAM (Generative Recursive reAsoning Models) introduces a stochastic framework where reasoning is no longer a single fixed path, but a distribution of possible trajectories. By injecting stochastic guidance into latent states and training via variational inference, GRAM solves complex puzzles like Sudoku-Extreme and ARC-AGI more effectively than deterministic SOTA models, while enabling parallel inference-time scaling.
Motivation: The Trap of Deterministic Latent States
Traditional Recursive Reasoning Models (RRMs) such as HRM or TRM operate by repeatedly applying the same Transformer block to a latent state, gradually "polishing" an answer. However, these models suffer from two fatal flaws:
- Irreversibility: If a deterministic update leads the model into a "reasoning dead end," it has no mechanism to backtrack or explore alternatives.
- Mode Collapse: In problems with multiple valid solutions (e.g., Graph Coloring), a deterministic model can only ever output one, failing to represent the true diversity of the solution space.
The authors' insight is profound: Reasoning should be generative. If we treat the internal "train of thought" as a stochastic process, the model can maintain multiple hypotheses and navigate complex constraint landscapes without getting stuck.
Methodology: Stochastic Guidance & Hierarchical Latent States
GRAM utilizes a nested loop structure to manage computational complexity:
- The Hierarchical State: The latent state is split into (high-level abstract reasoning) and (low-level fine-grained computation).
- Stochastic Transitions: Instead of a simple , GRAM uses: where is the stochastic guidance. This allows the model to "nudge" its reasoning path in different directions based on learned uncertainty.
Figure 1: The GRAM architecture showing the hierarchical update and the injection of stochastic guidance into the high-level state.
To train this, GRAM uses Amortized Variational Inference. During training, the model sees the target to learn a "perfect" trajectory (posterior); at test time, it relies on its learned "prior" to explore potential solutions.
Experiments: Superior Scaling and Multi-Solution Mastery
1. Breaking Scaling Bottlenecks
Deterministic models only scale by getting "deeper" (more iterations), which increases latency. GRAM introduces Width-Scaling: sampling trajectories in parallel.
- Result: 20 parallel samples at 16 iterations outperformed a deterministic TRM at 320 iterations. Parallelism bypasses the sequential bottleneck of depth.
Figure 2: Performance scaling on Sudoku-Extreme. GRAM (green/blue lines) scales effectively with both depth (x-axis) and width (N samples), significantly beating deterministic TRM (red).
2. Multi-Solution Tasks
In N-Queens and Graph Coloring, GRAM achieved nearly 90% coverage of all possible valid solutions, whereas deterministic baselines collapsed to ~30%. This proves that GRAM successfully internalizes the "manifold" of valid reasoning.
3. Unconditional Generation
Remarkably, GRAM can generate Sudoku boards from a blank grid with 99.05% validity. This demonstrates that the model doesn't just "solve" given puzzles but understands the underlying "grammar" of the constraints.
Critical Analysis & Future Outlook
Takeaway: This work bridges the gap between latent-state RRMs and generative models. By showing that "reasoning is sampling," it opens a new path for inference-time compute scaling that is parallelizable.
Limitations:
- Training Complexity: The deep supervision and sequential nature of recursion make it harder to scale to trillion-parameter foundation models compared to pure "next-token" prediction.
- Search Overhead: While parallel sampling is faster than sequential depth, it still requires significantly more total compute than a single-pass Transformer.
The Verdict: GRAM proves that "Thinking, Fast and Slow" can be implemented in a unified latent-variable framework. Future LLMs might not just predict the next token, but parallelize their "internal thoughts" through stochastic latent trajectories like GRAM.
Figure 3: Visualization of 50 sampled trajectories in latent space. Note how trajectories explore different regions to avoid local minima (yellow) and find the global optimum (dark blue).
