← Back to all articles
arXiv cs.CLSeptember 24, 2026

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

Excerpt

arXiv:2609.27373v1 Announce Type: new Abstract: Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant context, wh