LT2: Transforming Looped Depth into Linear-Time Intelligence
LT2: Linear-Time Looped Transformers
The paper introduces LT2 (Linear-Time Looped Transformers), a novel architecture family that replaces quadratic softmax attention in depth-recurrent (looped) models with subquadratic primitives like linear and sparse attention. By combining looping with these primitives, LT2 achieves high parameter efficiency and language modeling performance that matches or exceeds standard Looped Transformers while maintaining linear-time complexity and expanding effective receptive fields.
TL;DR
Looped Transformers are an elegant solution for parameter efficiency, but they have traditionally been trapped by the quadratic cost of attention. LT2 (Linear-Time Looped Transformers) breaks this cycle by replacing softmax attention with linear and sparse alternatives. The result is a model that is faster, more memory-efficient, and surprisingly more capable—matching the performance of 4B-parameter industry models with only 1.4B parameters.
The "Quadratic Loop" Problem
The industry is obsessed with scaling. While most scale by adding parameters, Looped Transformers (LT) scale by "reusing" parameters across multiple iterations (depth-wise recurrence). However, a fatal flaw remained: every loop re-calculated full attention.
If you have a sequence of 8k tokens and you loop 4 times, you aren't just doing 8k quadratic work once; you're doing it four times. This leads to a "decode cliff" where inference speed plummets and KV-cache memory explodes.
Methodology: The Synergy of Looping and Linear Mixers
The core insight of LT2 is that looping doesn't just save parameters—it actually "fixes" the inherent weaknesses of subquadratic mixers.
1. Rank-T Memory Refinement
Linear attention (like GDN or Mamba) typically updates its internal state with a rank-1 perturbation. Mathematically, this limits how much information can be "erased" or "replaced" in a single step. By looping times, LT2 effectively performs a Rank-T update, allowing the model to refine its internal "latent thoughts" far more precisely than a single-pass linear model.
2. Receptive Field Expansion
Sparse attention (like sliding windows) normally can't see beyond its local window . However, by looping times, the information travels further in each pass. LT2 expands the receptive field to , allowing a small, efficient window to eventually "see" the entire long-context sequence.
Figure: The LT2 hybrid strategies—interleaving mixers across depth and iterations.
The Pareto Frontier: Experiments and Results
The authors tested three main variants:
- LT2-Linear (GDN): Pure linear-time, high stability.
- LT2-Sparse (DSA): High precision, fixed context expansion.
- LT2-Hybrid (Full+GDN): The "Goldilocks" version.
Performance vs. Efficiency
The results are striking. In language modeling tasks, the LT2-Hybrid (Full+GDN) achieved a score of 62.89%, significantly higher than the standard Looped Transformer reference (59.27%).
Figure: LT2 models eliminate the "decode cliff," maintaining flat throughput as sequence length grows, whereas standard LTs eventually run out of memory (OOM).
Turning "Ouro" into "Ouro-Hybrid"
The paper doesn't just stop at training from scratch. They provide a distillation recipe to convert pre-trained full-attention looped models into LT2 variants. Their Ouro-hybrid-1.4B model matches the performance of industry-level 4B models while retaining the speed benefits of linear-time attention.
Critical Insight: The Attention Sink Mitigation
The authors identified a hidden pathology: Attention Sinks. In looped models, the tendency of attention to "dump" mass on the first token compounds across iterations, destabilizing the model. LT2 introduces a Gated Attention (SDPA output gate) which flattens these massive activations, proving essential for stability in recursive depth.
Conclusion
LT2 effectively argues that we don't need to choose between the parameter efficiency of looped models and the computational efficiency of linear attention. By combining them, we get a "latent reasoning" engine that can process massive contexts on a fraction of the hardware. This paves the way for a new generation of high-capability Small Language Models (SLMs) that can finally handle the "Infinite Context" dream efficiently.
For the full implementation and checkpoints, visit the LT2 GitHub Repository.
