← Back to all articles
arXiv cs.AIOctober 7, 2026

ReMem: Streaming Video Understanding With Long Context Retention

Excerpt

arXiv:2610.05940v1 Announce Type: cross Abstract: Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies in an attempt to adapt offline VLMs for streaming video understanding tasks. However, through our probing experiment, we identify