arXiv cs.AIAugust 18, 2026
KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving
Excerpt
arXiv:2608.15797v1 Announce Type: new Abstract: KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching the length limit. We characterize much of this loss as an information gapf caused by missing context, rather than a capability gap ca