arXiv cs.AIOctober 7, 2026
SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs
Excerpt
arXiv:2610.06413v1 Announce Type: cross Abstract: Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic