← Back to all articles
arXiv cs.LGOctober 2, 2026

Vision-language models for chest radiography do not always need the image

Excerpt

arXiv:2606.17710v3 Announce Type: replace-cross Abstract: Vision-language models that answer questions about chest radiographs are evaluated by their accuracy on labels derived from radiology reports. High benchmark accuracy is often interpreted as evidence that the model uses the image. A model that answers from the finding named in the question can score as well as a model that uses the radiograph. Keeping the question fixed, we audit eight open-weight systems by swapping in another patient's