ZEDA: Transforming Static MoE into Dynamic Transformers by Skipping Half the Experts
Post-Trained MoE Can Skip Half Experts via Self-Distillation
This paper introduces Zero-Expert Self-Distillation Adaptation (ZEDA), a low-cost framework that converts post-trained static Mixture-of-Experts (MoE) models into efficient dynamic ones. By injecting parameter-free "zero experts" and utilizing a two-stage self-distillation process, ZEDA enables models like Qwen3 and GLM-4 to skip over 50% of expert FLOPs while maintaining marginal accuracy loss and achieving roughly 1.2x end-to-end inference speedup.
TL;DR
Mixture-of-Experts (MoE) is the current standard for scaling LLMs, but static top-K routing often wastes computation on "easy" tokens. This paper presents ZEDA (Zero-Expert Self-Distillation Adaptation), a method to convert fully-trained, high-performance static MoEs into dynamic ones. By injecting "zero-computation" experts and using a specialized self-distillation pipeline, the authors prove you can skip 50% of expert processing with almost no loss in reasoning or coding ability, gaining a 20% real-world speedup.
Problem & Motivation: The Heavy Cost of "Static" Sparsity
Current MoE models (like Mixtral or DeepSeek-V3) use a fixed number of experts per token (e.g., top-2 or top-8). However, not every token requires the same amount of "brain power." A simple comma doesn't need 8 specialized experts; a complex math equation might.
The challenge? Converting a model that has already finished its expensive pre-training and SFT pipeline into a dynamic version is risky. Direct modifications usually break the delicate expert specializations, leading to catastrophic performance drops. ZEDA addresses this by making the transition "painless" via a teacher-student framework where the original model guides its own compression.
Methodology: Zero Experts and Group Balancing
ZEDA's magic lies in two components: Zero-Expert Injection and Group Auxiliary Loss.
1. Zero-Expert Injection
Instead of changing the MoE logic, ZEDA adds "Zero Experts" — parameterless modules that always return zero. By expanding the expert pool but keeping the top-K selection constant, the router can choose a "Zero Expert" to effectively "do nothing" for that slot, saving 100% of the FFN computation for that expert.
2. Group Auxiliary Loss (LGA)
Standard balancing losses try to use every expert equally. This is a disaster for post-trained models because it destroys the specialized routing learned during training. ZEDA introduces a Group-level strategy:
- Group E: Original normal experts.
- Group Z: New zero experts.
The loss only balances the ratio between Group E and Group Z, leaving the internal distribution of Group E untouched. This preserves the model’s established knowledge while forcing it to learn when it can afford to "skip" work.

Two-Stage Self-Distillation
To stabilize the new architecture, the model undergoes:
- SFT Stage: Learning to match the teacher's output distribution on fixed data.
- On-Policy Distillation (OPD): The model generates its own responses and is corrected by the teacher, closing the "distribution shift" gap that occurs during real-world inference.
Experiments & Results: Efficiency without Compromise
The researchers tested ZEDA on Qwen3-30B-A3B and GLM-4.7-Flash.
- Computational Savings: Specifically, ZEDA achieved a zero-expert activation ratio () of over 51.2%.
- Performance Preservation: On math (AIME), code (LiveCodeBench), and instruction following (IFEval), the ZEDA-adapted models remained highly competitive, often within 1% of the original static model.
- Inference Speed: Real-world tests using SGLang showed a ~1.20x speedup in both prefill and decode stages.

Deep Insight: What does the model skip?
An intriguing finding in the paper is that the model doesn't just skip "easy" tasks. Instead, zero-expert activation correlates with:
- Model Uncertainty: Tokens with high entropy (where the model is "confused") get more experts.
- Logp-Diff: Tokens where the student deviates from the teacher get more compute.
- Response Patterns: Structured segments like math and code expressions surprisingly allowed for higher skip rates after the initial "thinking" phase was stabilized.
Critical Analysis & Conclusion
ZEDA is a milestone for practical LLM deployment. It proves that the "last mile" of efficiency doesn't require retraining from scratch.
Limitations: The speedup benefit slightly decays as sequence length grows because the Attention mechanism (which is ) starts to dominate the total FLOPs, overshadowing the FFN savings. However, for most common context windows (up to 8k), the 20% gain is a significant win for high-throughput serving environments.
For developers and researchers, the takeaway is clear: Sparse models contain significant computational redundancy that can be harvested post-training through specialized self-distillation.
