arXiv cs.AIOctober 7, 2026
MOIRA: Mass-Oriented Indexing with Ragged Attention for Long-Context Decoding
Excerpt
arXiv:2610.04313v1 Announce Type: new Abstract: Long-context decoding is limited by memory bandwidth, because every output token reads the KV cache of every layer. Sparse decoding reduces this cost by reading only part of the KV cache. We observe that the number of pages a query needs varies widely across KV heads, layers and steps. Fixed budgets are simple, but they are sized for demanding cases and tuned per workload; adaptive budgets follow this variation more flexibly, but existing designs p