← Back to all articles
arXiv cs.AIAugust 18, 2026

MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

Excerpt

arXiv:2608.15299v1 Announce Type: cross Abstract: Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heterogeneity in layer-wise redundancy. We demonstrate that this uniformity is systematically suboptimal and propose MAPLE, a plug-and-play framework that reallocates the routed-expert budget heterogeneously across layers of any pretrained MoE LLM, without modifying weights or