Dynamic Short Convolutions: The "Locality Bias" Transformers Were Missing

Dynamic Short Convolutions Improve Transformers

2026-06-01
Oliver Sieberling, Bharat Runwal, Rameswar Panda, Yoon Kim
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Dynamic Short Convolutions (DSCs) as a scalable neural network primitive to improve Transformers by replacing static filters with input-dependent ones. Evaluated on models up to 2B parameters, the method achieves significant SOTA performance gains, notably a 1.60x compute advantage over standard Transformers when applied to all linear layers.

TL;DR

Researchers from MIT and IBM have introduced Dynamic Short Convolutions (DSC), a new architectural primitive that supercharges Transformers by allowing them to learn input-dependent local mixing. Unlike static convolutions, DSC filters adapt to every token. The result? A 1.60x compute advantage over standard Transformers and significant boosts in complex reasoning tasks like Fuzzy Recall.

The Motivation: Why Static Convolutions Aren't Enough

Transformers are global by nature, but language is often local. A phrase like "the old can swim" vs. "the old can opener" requires the model to compose meaning differently within a small window based on the final word.

While researchers have previously tried adding static short convolutions (fixed local filters) to models like Primer or Mamba, these fixed filters are "blind" to the specific nuances of the tokens they are processing. The authors argue that a truly powerful local primitive must be Dynamic—the filter itself should be a function of the content.

Methodology: Adaptive Filtering at Scale

The core idea is simple: instead of a fixed weight , use a weight generator that produces a filter based on the current hidden state .

1. Mathematical Intuition

Standard convolution:
Dynamic convolution:

This allows each token to decide how it wants to aggregate information from its immediate neighbors.

2. Solving the Parameter Explosion

Generating a full filter for every channel would double the model's parameters. To solve this, the authors use two clever tricks:

  • Low-Rank Parameterization: Factoring the weight generation through a low-rank bottleneck (e.g., Rank=16).
  • Head-wise Tying: Sharing the same dynamic filter across multiple channels (similar to Multi-Head Attention logic).

3. Hardware Efficiency (The Triton Secret)

Dynamic convolutions are usually memory-bound. To make this practical for training LLMs, the authors wrote custom Triton kernels that fuse the weight generation and the convolution on-chip. This prevents the large generated weights from ever being written to the slow Global Memory (HBM), keeping the overhead as low as 8%.

Model Architecture and Latency Figure 1: Comparison of Triton kernel latency vs. optimized baselines. The orange bars show our efficiency compared to existing CUDA-optimized static kernels.

Experiments: Dominating the Scaling Laws

The authors didn't just test on toy tasks; they applied DSC to dense Transformers, Mixture-of-Experts (MoE), and even Linear RNNs like Mamba-2.

Key Performance Metrics:

  • Compute Advantage: In a "compute-matched" setup, a Transformer with DSC achieves a lower loss than a vanilla Transformer that used 60% more compute.
  • Associative Recall: On tasks like "Fuzzy Recall" (retrieving values where keys have variable lengths), DSC models crushed the competition, highlighting the power of input-dependent local structure.

Scaling Law Improvements Figure 2: Scaling laws show that placing dynamic convolutions after every linear layer (right) provides a massive efficiency shift compared to standard Transformers.

Critical Analysis & Future Outlook

The beauty of this work lies in its universality. It’s not a replacement for Attention; it’s an enhancement to the linear layers ( or ) that already exist.

Why it works:

It gives the model a "cheap" way to handle local syntax before the "expensive" Attention mechanism has to deal with global semantics. It essentially acts as a highly sophisticated local feature extractor that updates at every step.

Limitations:

While the Triton kernels are fast, the "all-linear" variant still adds a ~22% throughput penalty. For inference-heavy production environments, further fusion (e.g., merging the convolution into the preceding MatMul) will be necessary to achieve true parity.

Conclusion

Dynamic Short Convolutions prove that we haven't reached the end of architectural innovation for LLMs. By adding a dash of "input-dependence" to local operations, we can build models that are not only more expressive but significantly more compute-efficient.

Find Similar Papers

Try Our Examples

  • Search for recent papers that integrate dynamic or input-dependent filtering mechanisms within the Transformer architecture to solve local dependency issues.
  • Which paper first introduced the concept of input-dependent dynamic filters for computer vision, and how does this paper adapt that theory for autoregressive language modeling?
  • Explore if dynamic short convolutions have been applied to multi-modal models or audio processing where local temporal structure is as critical as in text.
Contents
Dynamic Short Convolutions: The "Locality Bias" Transformers Were Missing
1. TL;DR
2. The Motivation: Why Static Convolutions Aren't Enough
3. Methodology: Adaptive Filtering at Scale
3.1. 1. Mathematical Intuition
3.2. 2. Solving the Parameter Explosion
3.3. 3. Hardware Efficiency (The Triton Secret)
4. Experiments: Dominating the Scaling Laws
4.1. Key Performance Metrics:
5. Critical Analysis & Future Outlook
5.1. Why it works:
5.2. Limitations:
6. Conclusion