← Back to all articles
arXiv cs.LGOctober 1, 2026

Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does

Excerpt

arXiv:2609.39892v1 Announce Type: new Abstract: Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises