← Back to all articles
arXiv cs.LGOctober 2, 2026

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

Excerpt

arXiv:2610.00418v1 Announce Type: new Abstract: Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce Comm