arXiv cs.AIOctober 7, 2026
Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs
Excerpt
arXiv:2610.05715v1 Announce Type: cross Abstract: Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitude estimation, and multi-view rotation reasoning. Across 12 VLMs, the gap is stark: the strongest models approach 100% accuracy on detection, yet even coarse magnitude estimation is near cha