← Back to all articles
arXiv cs.AIOctober 2, 2026

When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts

Excerpt

arXiv:2610.01243v1 Announce Type: cross Abstract: Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the returned image with human ratings the judge never sees and with random choice from the same candidates, and re