GRAM: Breaking the Deterministic Barrier in Recursive Reasoning
Generative Recursive Reasoning
Generative Recursive reAsoning Models (GRAM) is a novel framework that transforms deterministic recursive reasoning into a probabilistic generative process. By introducing stochastic latent transitions, GRAM achieves SOTA performance on structured reasoning benchmarks, notably reaching 97.0% on Sudoku-Extreme and outperforming prior recursive baselines like TRM and HRM.
TL;DR
Generative Recursive reAsoning Models (GRAM) represents a paradigm shift from deterministic to probabilistic latent reasoning. While previous recursive models followed a single, fixed path to a solution, GRAM utilizes stochastic latent transitions to explore multiple reasoning trajectories simultaneously. This allows for superior performance on complex puzzles (Sudoku-Extreme), multi-solution discovery (N-Queens), and even unconditional generation from empty inputs.
Problem & Motivation: The Deterministic Dead-End
Recursive Reasoning Models (RRMs) like HRM and TRM have shown that we can increase a model's "intelligence" by simply applying the same weights over and over—essentially thinking longer without adding parameters. However, these models were deterministic.
In complex reasoning, you often hit a "fork in the road." A deterministic model is forced to pick one path and stay on it. If that path is a local minimum, the model fails. Prior work lacked the "imagination" to maintain multiple hypotheses or backtrack through latent exploration. GRAM's core insight is that reasoning should be modeled as a stochastic process, not a straight line.
Methodology: High-Level Stochastic Guidance
GRAM's architecture is a hierarchical recycler. It splits the latent state into a high-level (h) and low-level (l) component.
- Inner Loop: The low-level state is refined deterministically to handle fine-grained computation.
- Outer Loop: The high-level state receives a Stochastic Guidance (SG) signal. Instead of a fixed update, the model samples:
This essentially "nudges" the reasoning trajectory, allowing the model to jump out of local minima.
Figure: The GRAM architecture showcasing the interaction between deterministic proposals () and stochastic guidance ().
Experiments & Results: Width vs. Depth
One of the most profound findings of GRAM is Inference-Time Scaling. Traditionally, more compute meant more depth (iterations). GRAM introduces width (parallel samples).
1. Scaling Success
The researchers found that sampling multiple trajectories in parallel (and picking the best via a Reward Model) is more efficient than just going deeper. GRAM with 20 parallel samples at 16 steps crushed TRM running at 320 steps.
Figure: Performance improvements on Sudoku-Extreme scaling with both depth (iterations) and width (N samples).
2. Multi-Solution Coverage
In N-Queens or Graph Coloring, there isn't just one "right" answer. Deterministic models collapse (mode collapse), capturing only ~36% of solutions. GRAM, through its stochastic nature, explores the solution space and achieves nearly 90% coverage.
3. Unconditional Generation
Remarkably, GRAM can generate valid Sudoku boards and MNIST digits starting from a blank grid. This proves it isn't just a solver—it's a true generative model that understands the underlying distribution of constraints.
Figure: Visualization of 50 sampled trajectories in the latent space. While TRM (deterministic) follows one path, GRAM explores the landscape to find the global optimum.
Critical Analysis & Conclusion
Takeaway
GRAM proves that probabilistic recursion is a fundamental design principle for future reasoning architectures. It decouples the "thinking" process from the token-generation process, allowing for compact, efficient models that can explore complex solution spaces.
Limitations
The primary bottleneck is the sequential nature of deep supervision. Training these models is memory-intensive and slow because you must backpropagate through many steps of the recursion, making it harder to scale to "Foundation Model" sizes compared to standard parallel Transformers.
Future Work
The next frontier is applying GRAM's principles to Multi-modal Large Language Models. Imagine an LLM that doesn't just predict the next token, but stochastically refines a "latent thought" until it reaches high confidence, scaling its "thinking time" dynamically based on task difficulty.
