arXiv cs.AIOctober 2, 2026
Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
Excerpt
arXiv:2610.01595v1 Announce Type: cross Abstract: Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $\tau_l$,