arXiv cs.AIOctober 7, 2026
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
Excerpt
arXiv:2606.24797v2 Announce Type: replace-cross Abstract: Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on Video Question Answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the faithfulness of predicted evidence supporting those answers remains insufficiently evaluated. This disconnect between answer generation and evidence verification motivates the construction of the Evidence-Gro