arXiv cs.AIOctober 7, 2026
What Keeps Vision-Language Models Looking at the Image?
Excerpt
arXiv:2607.12815v3 Announce Type: replace Abstract: When do vision-language models need direct access to the image while generating an answer? We study image dependence during answer generation by examining how the visual information needed for the current question becomes available in context. We intervene on direct image access while retaining previously computed states. Across real-image and synthetic tasks, we show that, depending on the generation process, direct access can continue to supp