arXiv cs.AIOctober 7, 2026
CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training
Excerpt
arXiv:2610.05744v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in