arXiv cs.LGOctober 2, 2026
Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization
Excerpt
arXiv:2610.00558v1 Announce Type: new Abstract: While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE