← Back to all articles
arXiv cs.AIOctober 7, 2026

Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding

Excerpt

arXiv:2610.06018v1 Announce Type: cross Abstract: Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current mo