← Back to all articles
arXiv cs.AIOctober 2, 2026

Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

Excerpt

arXiv:2610.01640v1 Announce Type: cross Abstract: Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone s