Oryx: Bridging the Gap Between Attention and Efficiency via Sequence-Axis Hybridization
Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
Oryx is a novel sequence-axis hybrid architecture that enables flexible switching between quadratic softmax attention and linear recurrent mixers (like Mamba-2 or Gated DeltaNet) within a single sequence. By tying over 90% of parameters and using shared key-value representations, Oryx achieves SOTA performance for hybrid models, outperforming pure baselines by 0.7+ points at the 1.4B scale on language modeling tasks.
TL;DR
Oryx is a breakthrough hybrid architecture that allows a Large Language Model to switch between Softmax Attention and Linear Recurrence (SSMs) on-the-fly within the same sequence. By sharing over 90% of weights and maintaining "compatible states," it offers the retrieval power of Transformers with the throughput of Mamba, outperforming both at the 1.4B scale.
Problem & Motivation: The Static Hybrid Bottleneck
In the current LLM landscape, we are forced to choose between the quadratic cost of Transformers (high quality, slow long-context) and the linear cost of SSMs/Linear Attention (lower quality/recall, extremely fast).
Existing solutions like Jamba or Zamba use "Inter-layer" hybrids (stacking different blocks). However, these are static. A model might need high-precision attention for a specific "needle" in a prompt but only requires fast linear recurrence for a reasoning chain. Oryx addresses this by allowing the model to change its "Mixer" mode at every chunk of tokens.
Methodology: The Oryx Shared Block
The technical core of Oryx is the Shared Key-Value Association. Despite their different mathematical formulations, both Attention and SSMs can be viewed as mechanisms that store and retrieve Key-Value pairs.
1. Unified State Updates
As shown in the architecture below, Oryx computes shared Key () and Value () representations. Even if only one mixer (e.g., Mamba-2) is used to produce the output for a chunk, the model simultaneously updates both the KV Cache and the Linear Recurrent State. This ensures that if the model switches to Attention in the next step, the KV Cache is "up to date."

2. Disjoint Queries
The authors found that while and can be shared, sharing the Query () projection hurts performance. Different mixers need unique "read-out" heads to extract information from the shared memory effectively.
3. Chunked Mixed-Mode Training
To make switching seamless, Oryx was trained by splitting sequences into 128-token chunks, with each chunk randomly assigned to a mixer mode. This "forces" the representations to remain compatible across transitions.
Experiments: Best of Both Worlds
Oryx was tested using Mamba-2 (Oryx-TM) and Gated DeltaNet (Oryx-TG) as linear mixers.
SOTA Performance
At 1.4B parameters, Oryx didn't just match baselines—it beat them. On average language modeling tasks, Oryx-TG reached 57.3% accuracy, significantly higher than the pure Gated DeltaNet (56.6%) or Transformer (55.6%) baselines.

Retrieval: The Ultimate Stress Test
The most impressive result lies in "Mixed Inference." By using the linear mode for the long context (prefill) and switching to Attention only for the final query (generation), Oryx achieved massive gains in Needle-in-a-Haystack (NIAH) tests—surpassing linear baselines by over 38 percentage points while maintaining much lower compute than a full Transformer.

Figure: The perplexity remains stable even after multiple switches, proving that the shared representations are truly compatible.
Critical Analysis & Conclusion
Oryx demonstrates that the "Intelligence vs. Efficiency" trade-off is not a zero-sum game.
Takeaways:
- Representational Compatibility: Attention and SSMs discover similar underlying features during training if forced through shared weights.
- Dynamic Compute: We can now imagine a "Thinking Model" that uses cheap linear modes for internal monologues and expensive attention for final output synthesis.
Limitations:
- Memory Overhead: During mixed-mode inference, the model must maintain both the KV Cache and the SSM state. While the KV Cache dominates at length, the SSM state adds a constant memory floor.
- Training Complexity: The chunked training strategy requires careful balancing of learning rates to remain stable across different optimization landscapes.
Oryx marks a pivotal shift toward Sequence-Axis Hybridization, providing a framework where efficiency and performance are allocated exactly where they are needed most.
