GRAM: Beyond Deterministic Logic with Probabilistic Recursive Reasoning
Generative Recursive Reasoning
This paper introduces Generative Recursive reAsoning Models (GRAM), a framework that transforms deterministic recursive reasoning into a probabilistic latent-variable generative process. By incorporating stochastic latent transitions, GRAM enables multi-hypothesis exploration and achieves SOTA performance on structured reasoning tasks like Sudoku-Extreme and ARC-AGI.
TL;DR
Generative Recursive reAsoning Models (GRAM) redefine internal model "thinking" not as a single fixed path, but as a probabilistic exploration of possibilities. By injecting stochasticity into recursive latent states, GRAM achieves superior accuracy on complex puzzles (Sudoku-Extreme, ARC-AGI) and enables "width-based" scaling—finding better answers by looking at multiple paths simultaneously rather than just thinking longer.
Background Positioning
In the quest for efficient reasoning, we've seen a move from "Chain-of-Thought" tokens to Recursive Reasoning Models (RRMs). RRMs use weight-sharing to refine a hidden state iteratively, effectively decoupling computational depth from parameter count. However, models like HRM and TRM are "stiff"—they are deterministic. If they start a reasoning path incorrectly, they stay on it. GRAM introduces stochastic guidance, turning these models from rigid calculators into generative explorers.
Problem & Motivation: The Trap of Deterministic Search
Why does a 10M parameter model fail where a human succeeds? Often, it's because the model gets stuck in a "local minimum" of logic. In tasks like Sudoku or N-Queens, one wrong early assumption can make the rest of the puzzle impossible.
Current RRMs follow a single attractor. If there are multiple valid solutions (like in Graph Coloring), a deterministic model can only ever see one. This mode collapse limits the model's robustness. The authors' insight is simple: if we treat reasoning as a stochastic process, we can sample multiple "trajectories" and use a reward model to pick the winner.
Methodology: High-Level Guidance, Low-Level Refinement
GRAM organizes computation into a hierarchical structure.
- Stochastic Latent Transitions: Instead of a fixed update , GRAM samples a perturbation from a learned Gaussian distribution.
- Hierarchical Loops:
- Inner Loop: A deterministic low-level component () handles fine-grained intermediate computation.
- Outer Loop: A high-level component () receives the stochastic guidance to steer the abstract reasoning direction.

The model is trained using Variational Inference. It learns to minimize a reconstruction loss (getting the right answer) while keeping its reasoning path close to a "prior" distribution, effectively learning a map of the "solution space."
Experiments: The Power of Width
The most striking result is Inference-Time Scaling. Traditionally, to get better results, you increase the number of iterations (Depth). GRAM allows you to increase the number of parallel samples (Width).
As shown in the charts below, GRAM scales significantly better than its predecessors. By sampling 20 different "thoughts" in parallel, it crushes the TRM baseline even when TRM is given 20 times more sequential depth.

In multi-solution tasks like N-Queens, GRAM maintains high accuracy regardless of the number of possible solutions, whereas deterministic models collapse as the problem becomes more "open-ended."
Deep Insight: Reasoning as Unconditional Generation
A fascinating byproduct of GRAM is its ability to generate content from scratch. By providing an "empty" input, the model can generate valid Sudoku boards or MNIST digits. This suggests that the model hasn't just memorized rules; it has internalized the "manifold" of valid structures.
Figure: GRAM progressively refines an image from noise into a coherent digit through recursive latent updates.
Critical Analysis & Conclusion
Takeaway: GRAM proves that the "Width" axis of inference compute—exploring multiple hypotheses in parallel—is a massive untapped resource for compact reasoning models.
Limitations:
- Training Complexity: The sequential nature of deep supervision makes it harder to parallelize training compared to standard Transformers.
- Latency vs. Compute: While parallel sampling reduces sequential latency, it increases the total VRAM/FLOP cost per inference.
Future Outlook: The ability to "steer" latent trajectories with stochastic guidance opens the door for Latent RL, where models could be fine-tuned to explore even more complex scientific or mathematical spaces without relying on discrete token-by-token generation.
