arXiv cs.CLSeptember 14, 2026
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Excerpt
arXiv:2609.13141v1 Announce Type: new Abstract: Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context units (tokens or blocks) for each query. Existing trainable methods usually use a lightweight selector to score context units, followed by hard Top-K selection that blocks gradients from the language modeling loss. Consequently, these methods commonly distill layer-wise dense attention distributions.