ZEDA: Transforming Static MoE into Dynamic Transformers by Skipping Half the Experts

Post-Trained MoE Can Skip Half Experts via Self-Distillation

2026-05-01
Xingtai Lv, Li Sheng, Kaiyan Zhang, Yichen You, Siyan Gao, Xueheng Luo, Yuxin Zuo, Yuchen Fan, Junlin Yang, Ganqu Cui, Bingning Wang, Fan Yang, Youbang Sun, Ning Ding, Bowen Zhou
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Zero-Expert Self-Distillation Adaptation (ZEDA), a low-cost framework that converts post-trained static Mixture-of-Experts (MoE) models into efficient dynamic ones. By injecting parameter-free "zero experts" and utilizing a two-stage self-distillation process, ZEDA enables models like Qwen3 and GLM-4 to skip over 50% of expert FLOPs while maintaining marginal accuracy loss and achieving roughly 1.2x end-to-end inference speedup.

TL;DR

Mixture-of-Experts (MoE) is the current standard for scaling LLMs, but static top-K routing often wastes computation on "easy" tokens. This paper presents ZEDA (Zero-Expert Self-Distillation Adaptation), a method to convert fully-trained, high-performance static MoEs into dynamic ones. By injecting "zero-computation" experts and using a specialized self-distillation pipeline, the authors prove you can skip 50% of expert processing with almost no loss in reasoning or coding ability, gaining a 20% real-world speedup.

Problem & Motivation: The Heavy Cost of "Static" Sparsity

Current MoE models (like Mixtral or DeepSeek-V3) use a fixed number of experts per token (e.g., top-2 or top-8). However, not every token requires the same amount of "brain power." A simple comma doesn't need 8 specialized experts; a complex math equation might.

The challenge? Converting a model that has already finished its expensive pre-training and SFT pipeline into a dynamic version is risky. Direct modifications usually break the delicate expert specializations, leading to catastrophic performance drops. ZEDA addresses this by making the transition "painless" via a teacher-student framework where the original model guides its own compression.

Methodology: Zero Experts and Group Balancing

ZEDA's magic lies in two components: Zero-Expert Injection and Group Auxiliary Loss.

1. Zero-Expert Injection

Instead of changing the MoE logic, ZEDA adds "Zero Experts" — parameterless modules that always return zero. By expanding the expert pool but keeping the top-K selection constant, the router can choose a "Zero Expert" to effectively "do nothing" for that slot, saving 100% of the FFN computation for that expert.

2. Group Auxiliary Loss (LGA)

Standard balancing losses try to use every expert equally. This is a disaster for post-trained models because it destroys the specialized routing learned during training. ZEDA introduces a Group-level strategy:

  • Group E: Original normal experts.
  • Group Z: New zero experts.

The loss only balances the ratio between Group E and Group Z, leaving the internal distribution of Group E untouched. This preserves the model’s established knowledge while forcing it to learn when it can afford to "skip" work.

Illustration of ZEDA

Two-Stage Self-Distillation

To stabilize the new architecture, the model undergoes:

  1. SFT Stage: Learning to match the teacher's output distribution on fixed data.
  2. On-Policy Distillation (OPD): The model generates its own responses and is corrected by the teacher, closing the "distribution shift" gap that occurs during real-world inference.

Experiments & Results: Efficiency without Compromise

The researchers tested ZEDA on Qwen3-30B-A3B and GLM-4.7-Flash.

  • Computational Savings: Specifically, ZEDA achieved a zero-expert activation ratio () of over 51.2%.
  • Performance Preservation: On math (AIME), code (LiveCodeBench), and instruction following (IFEval), the ZEDA-adapted models remained highly competitive, often within 1% of the original static model.
  • Inference Speed: Real-world tests using SGLang showed a ~1.20x speedup in both prefill and decode stages.

Experimental Results Comparison

Deep Insight: What does the model skip?

An intriguing finding in the paper is that the model doesn't just skip "easy" tasks. Instead, zero-expert activation correlates with:

  • Model Uncertainty: Tokens with high entropy (where the model is "confused") get more experts.
  • Logp-Diff: Tokens where the student deviates from the teacher get more compute.
  • Response Patterns: Structured segments like math and code expressions surprisingly allowed for higher skip rates after the initial "thinking" phase was stabilized.

Critical Analysis & Conclusion

ZEDA is a milestone for practical LLM deployment. It proves that the "last mile" of efficiency doesn't require retraining from scratch.

Limitations: The speedup benefit slightly decays as sequence length grows because the Attention mechanism (which is ) starts to dominate the total FLOPs, overshadowing the FFN savings. However, for most common context windows (up to 8k), the 20% gain is a significant win for high-throughput serving environments.

For developers and researchers, the takeaway is clear: Sparse models contain significant computational redundancy that can be harvested post-training through specialized self-distillation.

Find Similar Papers

Try Our Examples

  • Search for recent papers dealing with dynamic expert activation or expert skipping for large-scale Mixture-of-Experts models beyond the pre-training phase.
  • Which paper first introduced the concept of "zero-computation experts" or "null experts," and how does ZEDA's group-level balancing differ from the original implementation?
  • Investigate the application of on-policy distillation (OPD) for improving the inference efficiency of transformer-based architectures in domains like Computer Vision or Robotics.
Contents
ZEDA: Transforming Static MoE into Dynamic Transformers by Skipping Half the Experts
1. TL;DR
2. Problem & Motivation: The Heavy Cost of "Static" Sparsity
3. Methodology: Zero Experts and Group Balancing
3.1. 1. Zero-Expert Injection
3.2. 2. Group Auxiliary Loss (LGA)
4. Two-Stage Self-Distillation
5. Experiments & Results: Efficiency without Compromise
6. Deep Insight: What does the model skip?
7. Critical Analysis & Conclusion