← Back to all articles
arXiv cs.CLOctober 7, 2026

Hybrid Latent Attention for Looped Language Models

Excerpt

arXiv:2610.07940v1 Announce Type: new Abstract: Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a