arXiv cs.CLSeptember 14, 2026
Expert-Space Exploration in MoE Reinforcement Learning
Excerpt
arXiv:2609.13058v1 Announce Type: new Abstract: Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the sparse computation paths that induce output distributions, expert selection offers an additional source of rollout diversity. Through