arXiv cs.AIAugust 18, 2026
Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
Excerpt
arXiv:2608.16805v1 Announce Type: cross Abstract: Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neither reveals the transfer. This study formalizes this blind spot as Dense Same-Class Attribute Misbinding (DSCAM) and presents Insta