arXiv cs.AIOctober 2, 2026
CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
Excerpt
arXiv:2610.01710v1 Announce Type: new Abstract: Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region