← Back to all articles
arXiv cs.LGOctober 1, 2026

Switching Linear Attention

Excerpt

arXiv:2609.39034v1 Announce Type: new Abstract: Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expres