← Back to all articles
arXiv cs.LGOctober 2, 2026

Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

Excerpt

arXiv:2610.00024v1 Announce Type: cross Abstract: Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on