Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

2026-05-01
Ali Hatamizadeh, Yejin Choi, Jan Kautz
Summary
Problem
Method
Results
Takeaways
Abstract

Gated DeltaNet-2 is a novel linear recurrent attention architecture that decouples the memory "erase" and "write" operations through channel-wise gating. It generalizes existing models like KDA and Gated DeltaNet, achieving State-of-the-Art (SOTA) performance on long-context retrieval and language modeling benchmarks compared to Mamba-2/3 and other linear RNNs.

TL;DR

Linear recurrent architectures are the primary contenders to replace Transformers for long-context tasks, but they suffer from "memory interference." Gated DeltaNet-2 solves this by decoupling the memory update into two independent channel-wise gates: an erase gate (what to forget from the keys) and a write gate (what to commit to the values). This simple yet profound architectural shift allows it to dominate long-context retrieval benchmarks while maintaining superior efficiency.

Problem & Motivation: The Tug-of-War in Recurrent States

In a standard Linear Attention or State Space Model (SSM), the model must compress an entire sequence into a fixed-size matrix . The "Delta Rule" was introduced to improve this by subtracting old associations before writing new ones (i.e., overwriting).

However, previous SOTA models like Kimi Delta Attention (KDA) and Gated DeltaNet used a single scalar value to control this process. This created a "tug-of-war": the same coefficient determined how much of the old memory was erased and how much of the new value was added.

The authors' core insight is that erasing is a key-side operation (which addresses to clear?), while writing is a value-side operation (what information to store?). Tying them together is an unnecessary modeling restriction that limits the model's ability to handle complex associations.

Methodology: Gated Delta Rule-2

The authors propose a more expressive update mechanism, Gated Delta Rule-2, which uses channel-wise vectors instead of scalars:

  1. Selective Erase (): A gate that weights key coordinates used to read and remove old content.
  2. Selective Write (): A gate that weights the value coordinates being inserted into the state.

The state update becomes:

Overall Architecture

The Fast-Weight Perspective

From a "Fast-Weight Programmer" viewpoint, this update is essentially an online gradient step on a local regression loss. While Mamba-2/3 focus on correlation writes into a decayed state, Gated DeltaNet-2 performs a residual delta edit, which is inherently more precise for tasks involving specific fact retrieval.

Efficient Parallel Training

To make this trainable on modern GPUs, the authors derived a Gate-Aware Chunkwise Algorithm. By absorbing channel-wise decay into asymmetric erase factors, they utilize the "WY" form of Householder products. This allows the model to process chunks of tokens in parallel, achieving complexity with high hardware utilization via fused Triton kernels.

Experiments & Results

The model was tested at a 1.3B parameter scale across various settings: Recurrent-only and Hybrid (Recurrent + Sliding Window Attention).

1. Long-Context Power: RULER Benchmark

The most striking result is in the Multi-Key Needle-In-A-Haystack (MK-NIAH) test. In scenarios where a fixed-size state must separate multiple competing associations, Gated DeltaNet-2 significantly outperforms Mamba-2 and Mamba-3.

Experimental Results Table

2. Efficiency

Despite the added complexity of channel-wise gating, the training throughput remains remarkably flat. In an H100 comparison, Gated DeltaNet-2 shows only a marginal drop in speed compared to simpler linear models, while remaining significantly faster than standard Transformers at long sequence lengths.

Training Throughput

Critical Analysis & Conclusion

Takeaway

Gated DeltaNet-2 proves that the "Memory Bottleneck" of recurrent models is often a software/modeling bottleneck rather than a hardware one. By refining the how of memory editing (decoupling erase/write), we can push linear models much closer to the retrieval capabilities of the KV-cache.

Limitations & Future Work

While Gated DeltaNet-2 is superior in retrieval, the authors note that it still benefits from a Sliding Window Attention (SWA) hybrid setup for local evidence aggregation (e.g., in datasets like DROP). Future research may explore if even more complex "erase" mechanisms—perhaps multi-rank updates—could eventually eliminate the need for any attention windows entirely.


For more details, check out the official repository.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize or extend the "delta rule" within the context of linear attention or fast-weight programmers for large language models.
  • What are the theoretical origins of the "WY representation" for Householder matrices, and how has its application evolved in parallelizing recurrent neural networks?
  • Investigate the current SOTA methods for "multi-key retrieval" in long-context benchmarks (like RULER) that specifically compare hybrid SSM-Transformer architectures.
Contents
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
1. TL;DR
2. Problem & Motivation: The Tug-of-War in Recurrent States
3. Methodology: Gated Delta Rule-2
3.1. The Fast-Weight Perspective
3.2. Efficient Parallel Training
4. Experiments & Results
4.1. 1. Long-Context Power: RULER Benchmark
4.2. 2. Efficiency
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work