← Back to all articles
arXiv cs.AIOctober 2, 2026

Geometric Similarity in VLM Low-Level Vision Representations

Excerpt

arXiv:2610.00848v1 Announce Type: cross Abstract: Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a