arXiv cs.AIOctober 2, 2026
A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
Excerpt
arXiv:2610.00952v1 Announce Type: cross Abstract: Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($\pi$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choice