← Back to all articles
arXiv cs.LGOctober 2, 2026

Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization

Excerpt

arXiv:2610.00558v1 Announce Type: new Abstract: While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE