← Back to all articles
arXiv cs.AIOctober 7, 2026

MOIRA: Mass-Oriented Indexing with Ragged Attention for Long-Context Decoding

Excerpt

arXiv:2610.04313v1 Announce Type: new Abstract: Long-context decoding is limited by memory bandwidth, because every output token reads the KV cache of every layer. Sparse decoding reduces this cost by reading only part of the KV cache. We observe that the number of pages a query needs varies widely across KV heads, layers and steps. Fixed budgets are simple, but they are sized for demanding cases and tuned per workload; adaptive budgets follow this variation more flexibly, but existing designs p