arXiv cs.AIOctober 2, 2026
Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
Excerpt
arXiv:2610.02021v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (C