UniPool: Breaking the Linear Scaling Myth of MoE Expert Parameters
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
The paper introduces UniPool, a novel Mixture-of-Experts (MoE) architecture that replaces layer-specific expert sets with a globally shared expert pool. This unified pool, managed by per-layer independent routers and a new pool-level auxiliary loss, achieves state-of-the-art results across five LLaMA-based scales (up to 978M parameters).
TL;DR
UniPool is a revolutionary Mixture-of-Experts (MoE) architecture that ditches the traditional "one set of experts per layer" rule in favor of a Globally Shared Expert Pool. By allowing all layers to access the same resource pool, it achieves better performance with up to 60% fewer expert parameters, effectively decoupling model depth from parameter explosion.
Problem: The Hidden Waste in Modern MoE
Current MoE scaling is governed by a rigid law: if you add more layers, you must add more experts. This implies that every layer requires its own private set of "specialists." However, the authors' "routing probe" reveals a startling truth: in production models like Qwen and DeepSeek, if you randomize the routing in deep layers, the accuracy only drops by a measly 1.0–1.6 points.
This proves that deep-layer experts are largely redundant—each layer is essentially spending its budget to rediscover the same transformations.
Methodology: The Global Pool Strategy
UniPool transforms expert capacity from a "local property" to a "global architectural budget."
1. Global Expert Pool
Instead of layers having experts, UniPool has one pool of experts. This allows for expert reuse: a specific expert can now be triggered by Layer 2 and then again by Layer 20, providing a much higher return on the parameter "investment."
2. Pool-Level Auxiliary Loss
Standard load-balancing losses are per-layer. In a shared pool, this is counter-productive because forcing every layer to use every expert prevents specialization. UniPool uses a Pool-Level Loss that ensures every expert in the global pool is used by someone, but allows individual layers the freedom to focus on specific subsets.
3. NormRouter: Stability at Scale
Since hidden states change scale as tokens go deeper into the network, a standard Softmax router behaves inconsistently. UniPool employs NormRouter, which uses L2-normalization to ensure that the routing "sharpness" remains stable whether a token is in the 1st layer or the 48th layer.

Experiments: Doing More with Less
The authors tested UniPool across five scales using the LLaMA architecture. The results were consistent: UniPool didn't just beat vanilla MoE; it crushed it in terms of parameter efficiency.
- Efficiency Milestone: An 830M parameter model using only 41.6% of the vanilla expert budget still performed better than the standard MoE baseline.
- Scaling Insight: As models get deeper, the shared pool becomes even more effective. This suggests that "pool size" is a new hyperparameter that can grow sublinearly with depth.

Critical Analysis: Why This Matters
The most profound takeaway is that the routing decisions in UniPool are "load-bearing." In vanilla MoE, experts are so similar that the router's choice barely matters. In UniPool, because experts accumulate gradients from multiple layers, they become highly specialized. Randomizing a UniPool router causes a 4.1 point drop (compared to ~1.3 in vanilla), proving that the neurons are actually doing unique, specialized work.
Limitations
While the parameter efficiency is undeniable, the paper notes that routing into a larger global pool might introduce minor overheads in cross-layer statistics or token-dispatch efficiency in highly distributed settings (Expert Parallelism).
Conclusion
UniPool proves that we don't need a linear increase in parameters to gain model depth. By treating experts as a shared global utility, we can build deeper, smarter models that are significantly leaner than the current SOTA.
Figure: UniPool (bottom) achieves balanced global utilization while maintaining layer-specific routing signatures, unlike the collapsed distribution seen in shared-pool models using standard losses (top).
