← Back to all articles
arXiv cs.AIOctober 7, 2026

AnyGroundBench: A Multi-Domain Adaptation Benchmark for Video Grounding in VLMs

Excerpt

arXiv:2607.02269v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have shown strong performance in Spatio-Temporal Video Grounding (STVG), yet they are still evaluated mostly in a zero-shot manner on general-purpose benchmarks of everyday scenes. This creates a critical disconnect from real-world applications in specialized domains, where models inevitably encounter rare visual or textual concepts. Since exhaustive pre-training across infinite data distributions is infeasib