arXiv cs.AIOctober 7, 2026
Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents
Excerpt
arXiv:2610.05833v1 Announce Type: new Abstract: Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requ