LT2: Transforming Looped Depth into Linear-Time Intelligence

LT2: Linear-Time Looped Transformers

2026-05-01
Chunyuan Deng, Yizhe Zhang, Rui-Jie Zhu, Yuanyuan Xu, Jiarui Liu, T. S. Eugene Ng, Hanjie Chen
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces LT2 (Linear-Time Looped Transformers), a novel architecture family that replaces quadratic softmax attention in depth-recurrent (looped) models with subquadratic primitives like linear and sparse attention. By combining looping with these primitives, LT2 achieves high parameter efficiency and language modeling performance that matches or exceeds standard Looped Transformers while maintaining linear-time complexity and expanding effective receptive fields.

TL;DR

Looped Transformers are an elegant solution for parameter efficiency, but they have traditionally been trapped by the quadratic cost of attention. LT2 (Linear-Time Looped Transformers) breaks this cycle by replacing softmax attention with linear and sparse alternatives. The result is a model that is faster, more memory-efficient, and surprisingly more capable—matching the performance of 4B-parameter industry models with only 1.4B parameters.

The "Quadratic Loop" Problem

The industry is obsessed with scaling. While most scale by adding parameters, Looped Transformers (LT) scale by "reusing" parameters across multiple iterations (depth-wise recurrence). However, a fatal flaw remained: every loop re-calculated full attention.

If you have a sequence of 8k tokens and you loop 4 times, you aren't just doing 8k quadratic work once; you're doing it four times. This leads to a "decode cliff" where inference speed plummets and KV-cache memory explodes.

Methodology: The Synergy of Looping and Linear Mixers

The core insight of LT2 is that looping doesn't just save parameters—it actually "fixes" the inherent weaknesses of subquadratic mixers.

1. Rank-T Memory Refinement

Linear attention (like GDN or Mamba) typically updates its internal state with a rank-1 perturbation. Mathematically, this limits how much information can be "erased" or "replaced" in a single step. By looping times, LT2 effectively performs a Rank-T update, allowing the model to refine its internal "latent thoughts" far more precisely than a single-pass linear model.

2. Receptive Field Expansion

Sparse attention (like sliding windows) normally can't see beyond its local window . However, by looping times, the information travels further in each pass. LT2 expands the receptive field to , allowing a small, efficient window to eventually "see" the entire long-context sequence.

Model Architecture and Hybrid Patterns Figure: The LT2 hybrid strategies—interleaving mixers across depth and iterations.

The Pareto Frontier: Experiments and Results

The authors tested three main variants:

  • LT2-Linear (GDN): Pure linear-time, high stability.
  • LT2-Sparse (DSA): High precision, fixed context expansion.
  • LT2-Hybrid (Full+GDN): The "Goldilocks" version.

Performance vs. Efficiency

The results are striking. In language modeling tasks, the LT2-Hybrid (Full+GDN) achieved a score of 62.89%, significantly higher than the standard Looped Transformer reference (59.27%).

Efficiency Comparison Figure: LT2 models eliminate the "decode cliff," maintaining flat throughput as sequence length grows, whereas standard LTs eventually run out of memory (OOM).

Turning "Ouro" into "Ouro-Hybrid"

The paper doesn't just stop at training from scratch. They provide a distillation recipe to convert pre-trained full-attention looped models into LT2 variants. Their Ouro-hybrid-1.4B model matches the performance of industry-level 4B models while retaining the speed benefits of linear-time attention.

Critical Insight: The Attention Sink Mitigation

The authors identified a hidden pathology: Attention Sinks. In looped models, the tendency of attention to "dump" mass on the first token compounds across iterations, destabilizing the model. LT2 introduces a Gated Attention (SDPA output gate) which flattens these massive activations, proving essential for stability in recursive depth.

Conclusion

LT2 effectively argues that we don't need to choose between the parameter efficiency of looped models and the computational efficiency of linear attention. By combining them, we get a "latent reasoning" engine that can process massive contexts on a fraction of the hardware. This paves the way for a new generation of high-capability Small Language Models (SLMs) that can finally handle the "Infinite Context" dream efficiently.


For the full implementation and checkpoints, visit the LT2 GitHub Repository.

Find Similar Papers

Try Our Examples

  • Search for recent papers that explore the synergy between depth-wise weight sharing (Universal Transformers) and State Space Models (SSMs) or Linear Attention.
  • What are the primary theoretical limitations of rank-1 recurrent updates identified in prior work, and how does LT2 specifically address the Sn word problem through looping?
  • Find research evaluating the application of Looped Transformers or LT2-like hybrid architectures in multimodal (Vision/Audio) generation tasks to improve parameter efficiency.
Contents
LT2: Transforming Looped Depth into Linear-Time Intelligence
1. TL;DR
2. The "Quadratic Loop" Problem
3. Methodology: The Synergy of Looping and Linear Mixers
3.1. 1. Rank-T Memory Refinement
3.2. 2. Receptive Field Expansion
4. The Pareto Frontier: Experiments and Results
4.1. Performance vs. Efficiency
4.2. Turning "Ouro" into "Ouro-Hybrid"
5. Critical Insight: The Attention Sink Mitigation
6. Conclusion