← Back to all articles
arXiv cs.AIOctober 2, 2026

Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

Excerpt

arXiv:2610.02021v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (C