arXiv cs.AIOctober 7, 2026
LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention
Excerpt
arXiv:2610.04635v1 Announce Type: new Abstract: Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We introduce LatentIndex, which extends the latent-sharing principle of Multi-head Latent Attention across indexer layers. Each layer group constructs a shared latent cache from its first layer's hi