arXiv cs.LGOctober 7, 2026
Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
Excerpt
arXiv:2512.23073v2 Announce Type: replace Abstract: Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we