← Back to all articles
arXiv cs.CLSeptember 24, 2026

What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit

Excerpt

arXiv:2609.27408v1 Announce Type: cross Abstract: Benchmarks for vision-language models offer their answer choices in some convention: a letter, a color name, a pixel coordinate. That convention is treated as neutral. We find it is not, and that the limits a benchmark reports can belong to the readout rather than to the model. On 200 COCO photographs, Qwen3-VL-4B picks the correct one of nine locations for a named object 68.5% of the time when the locations are given in English and 20.0% when th