← Back to all articles
arXiv cs.LGOctober 2, 2026

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

Excerpt

arXiv:2610.01153v1 Announce Type: new Abstract: Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield li