arXiv cs.LGOctober 7, 2026
Distributionally Robust Mixture-of-Experts Training
Excerpt
arXiv:2610.07207v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss ro