arXiv cs.AIOctober 7, 2026
A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models
Excerpt
arXiv:2610.05413v1 Announce Type: cross Abstract: Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experiment