arXiv cs.CLSeptember 22, 2026
Block-Sparse Attention with Semantic-Geometric Decoupled Routing
Excerpt
arXiv:2609.22884v1 Announce Type: new Abstract: Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly alternative by routing each query block to a small set of relevant key blocks, yet accurate training-free block routing remains difficult. Existing routers often pool post-RoPE token representations, which entangles semantic agg