Oryx: Bridging the Gap Between Attention and Efficiency via Sequence-Axis Hybridization

Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations

2026-05-01
Kevin Y. Li, Asher Trockman, Ananda Theertha Suresh, Ziteng Sun
Summary
Problem
Method
Results
Takeaways
Abstract

Oryx is a novel sequence-axis hybrid architecture that enables flexible switching between quadratic softmax attention and linear recurrent mixers (like Mamba-2 or Gated DeltaNet) within a single sequence. By tying over 90% of parameters and using shared key-value representations, Oryx achieves SOTA performance for hybrid models, outperforming pure baselines by 0.7+ points at the 1.4B scale on language modeling tasks.

TL;DR

Oryx is a breakthrough hybrid architecture that allows a Large Language Model to switch between Softmax Attention and Linear Recurrence (SSMs) on-the-fly within the same sequence. By sharing over 90% of weights and maintaining "compatible states," it offers the retrieval power of Transformers with the throughput of Mamba, outperforming both at the 1.4B scale.

Problem & Motivation: The Static Hybrid Bottleneck

In the current LLM landscape, we are forced to choose between the quadratic cost of Transformers (high quality, slow long-context) and the linear cost of SSMs/Linear Attention (lower quality/recall, extremely fast).

Existing solutions like Jamba or Zamba use "Inter-layer" hybrids (stacking different blocks). However, these are static. A model might need high-precision attention for a specific "needle" in a prompt but only requires fast linear recurrence for a reasoning chain. Oryx addresses this by allowing the model to change its "Mixer" mode at every chunk of tokens.

Methodology: The Oryx Shared Block

The technical core of Oryx is the Shared Key-Value Association. Despite their different mathematical formulations, both Attention and SSMs can be viewed as mechanisms that store and retrieve Key-Value pairs.

1. Unified State Updates

As shown in the architecture below, Oryx computes shared Key () and Value () representations. Even if only one mixer (e.g., Mamba-2) is used to produce the output for a chunk, the model simultaneously updates both the KV Cache and the Linear Recurrent State. This ensures that if the model switches to Attention in the next step, the KV Cache is "up to date."

Oryx Shared Block Architecture

2. Disjoint Queries

The authors found that while and can be shared, sharing the Query () projection hurts performance. Different mixers need unique "read-out" heads to extract information from the shared memory effectively.

3. Chunked Mixed-Mode Training

To make switching seamless, Oryx was trained by splitting sequences into 128-token chunks, with each chunk randomly assigned to a mixer mode. This "forces" the representations to remain compatible across transitions.

Experiments: Best of Both Worlds

Oryx was tested using Mamba-2 (Oryx-TM) and Gated DeltaNet (Oryx-TG) as linear mixers.

SOTA Performance

At 1.4B parameters, Oryx didn't just match baselines—it beat them. On average language modeling tasks, Oryx-TG reached 57.3% accuracy, significantly higher than the pure Gated DeltaNet (56.6%) or Transformer (55.6%) baselines.

Performance Comparison Table

Retrieval: The Ultimate Stress Test

The most impressive result lies in "Mixed Inference." By using the linear mode for the long context (prefill) and switching to Attention only for the final query (generation), Oryx achieved massive gains in Needle-in-a-Haystack (NIAH) tests—surpassing linear baselines by over 38 percentage points while maintaining much lower compute than a full Transformer.

Perplexity under Mode-Switching

Figure: The perplexity remains stable even after multiple switches, proving that the shared representations are truly compatible.

Critical Analysis & Conclusion

Oryx demonstrates that the "Intelligence vs. Efficiency" trade-off is not a zero-sum game.

Takeaways:

  • Representational Compatibility: Attention and SSMs discover similar underlying features during training if forced through shared weights.
  • Dynamic Compute: We can now imagine a "Thinking Model" that uses cheap linear modes for internal monologues and expensive attention for final output synthesis.

Limitations:

  • Memory Overhead: During mixed-mode inference, the model must maintain both the KV Cache and the SSM state. While the KV Cache dominates at length, the SSM state adds a constant memory floor.
  • Training Complexity: The chunked training strategy requires careful balancing of learning rates to remain stable across different optimization landscapes.

Oryx marks a pivotal shift toward Sequence-Axis Hybridization, providing a framework where efficiency and performance are allocated exactly where they are needed most.

Find Similar Papers

Try Our Examples

  • Search for recent papers published after 2024 that implement dynamic routing between attention and state-space models at the token or chunk level.
  • Which study first demonstrated that initializing State Space Models to mimic Attention (Mimetic Initialization) improves recall, and how does Oryx's shared weight approach qualitatively differ from this?
  • Explore research applying sequence-axis hybridization or dynamic compute switching to multimodal models or video diffusion tasks.
Contents
Oryx: Bridging the Gap Between Attention and Efficiency via Sequence-Axis Hybridization
1. TL;DR
2. Problem & Motivation: The Static Hybrid Bottleneck
3. Methodology: The Oryx Shared Block
3.1. 1. Unified State Updates
3.2. 2. Disjoint Queries
3.3. 3. Chunked Mixed-Mode Training
4. Experiments: Best of Both Worlds
4.1. SOTA Performance
4.2. Retrieval: The Ultimate Stress Test
5. Critical Analysis & Conclusion