arXiv cs.AIOctober 7, 2026
Enhancing Long-Video VLM Embeddings with Query-Aware Streaming Latent Reasoning
Excerpt
arXiv:2610.04864v1 Announce Type: cross Abstract: Long-video embedding requires capturing sparse query-relevant evidence under a limited visual-token budget. Uniform sampling can miss brief events in videos spanning minutes or hours, whereas encoding more frames in a single context increases memory and computation. We introduce \textbf{Query-Aware Streaming Latent Reasoning} (QASLR), a post-training framework that accumulates evidence across clips while keeping the embedding size fixed. QASLR se