← Back to all articles
arXiv cs.AIOctober 7, 2026

More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding

Excerpt

arXiv:2610.04753v2 Announce Type: cross Abstract: Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate infe