arXiv cs.LGOctober 1, 2026
MoEless: Efficient MoE LLM Serving with Serverless Experts
Excerpt
arXiv:2603.06350v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incur