← Back to all articles
arXiv cs.AIOctober 7, 2026

CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs

Excerpt

arXiv:2610.06399v1 Announce Type: new Abstract: Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that