← Back to all articles
arXiv cs.CLSeptember 24, 2026

DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video

Excerpt

arXiv:2609.27470v1 Announce Type: cross Abstract: Recent video-language models increasingly adopt hybrid architectures that interleave linear and full attention layers for efficient long-context processing. While the recurrent state of linear attention remains fixed in size, the KV cache of full attention continues to grow with the video stream, making eviction necessary under a bounded memory budget. The key challenge in streaming is that eviction must occur before the question arrives, so what