← Back to all articles
arXiv cs.AIOctober 7, 2026

What Keeps Vision-Language Models Looking at the Image?

Excerpt

arXiv:2607.12815v3 Announce Type: replace Abstract: When do vision-language models need direct access to the image while generating an answer? We study image dependence during answer generation by examining how the visual information needed for the current question becomes available in context. We intervene on direct image access while retaining previously computed states. Across real-image and synthetic tasks, we show that, depending on the generation process, direct access can continue to supp