arXiv cs.LGOctober 7, 2026
Structuring MoE Expert Selection for Agentic Reinforcement Learning
Excerpt
arXiv:2610.07332v1 Announce Type: new Abstract: Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert rout