← Back to all articles
arXiv cs.AIOctober 2, 2026

A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

Excerpt

arXiv:2610.00952v1 Announce Type: cross Abstract: Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($\pi$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choice