← Back to all articles
arXiv cs.LGOctober 2, 2026

Decoding Looped Transformers Better for (Almost) Free

Excerpt

arXiv:2610.02185v1 Announce Type: new Abstract: Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contras