← Back to all articles
arXiv cs.CLSeptember 9, 2026

NEST: Narrative Event Structures in Time for Long Video Understanding

Excerpt

arXiv:2606.19706v2 Announce Type: replace-cross Abstract: Recent progress in vision-language models has enabled processing of increasingly long video sequences, but handling extended token streams does not translate to understanding complex narrative structure in long videos. Existing long-video benchmarks focus on needle-in-a-haystack retrieval rather than evaluating how low-level actions form events, interact across time, and drive narratives, for example whether a model can connect an early j