← Back to all articles
arXiv cs.AIOctober 2, 2026

CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

Excerpt

arXiv:2610.01710v1 Announce Type: new Abstract: Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region