← Back to all articles
arXiv cs.AIAugust 18, 2026

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

Excerpt

arXiv:2608.16805v1 Announce Type: cross Abstract: Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neither reveals the transfer. This study formalizes this blind spot as Dense Same-Class Attribute Misbinding (DSCAM) and presents Insta