arXiv cs.LGOctober 1, 2026
ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control
Excerpt
arXiv:2609.39137v1 Announce Type: new Abstract: Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However, further scaling the parameter count requires increasingly sparse routing, where expert load imbalance becomes more severe. This imbalance reduces parameter utilization and training efficiency, and can undermine training stability, becoming a bottleneck to reliable scaling. In this work, we unify two