arXiv cs.AIOctober 7, 2026
Localization Lens for Improving Medical Vision-Language Models
Excerpt
arXiv:2610.04502v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) have demonstrated strong capabilities in clinical tasks. However, they often struggle to understand anatomical structures and spatial positioning, which are crucial for medical reasoning. To address this, we propose a localization-aware enhancement to the Med-VLM pipeline, introducing improvements at three levels: data,architecture, and alignment. First, we introduce localization lens, a set of expert-val