arXiv cs.LGOctober 1, 2026
Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does
Excerpt
arXiv:2609.39892v1 Announce Type: new Abstract: Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises