arXiv cs.LGOctober 2, 2026
MoRA: MoE Pruning via Router Bias Learning and Expert Approximation
Excerpt
arXiv:2610.00367v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on eff