← Back to all articles
arXiv cs.LGOctober 2, 2026

EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction

Excerpt

arXiv:2610.00412v1 Announce Type: new Abstract: KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieves strong compression quality at the cost of additional forward passes. Learned approximations reduce this cost but require model-specific training. We analyze how KVzip identifies important cached information and show ho