arXiv cs.AIOctober 7, 2026
Trinity: Self-Evolving Vision-Language Models with a Self-Verifier
Excerpt
arXiv:2610.04469v1 Announce Type: new Abstract: Self-evolving vision-language models (VLMs), a form of self-improvement in which a model generates its own training data from unlabeled images, are a promising route toward agents that expand their reasoning capability in an unsupervised manner, without relying on ever-larger annotation budgets. Existing methods pair a Questioner that proposes problems with a Solver that answers them, but reward both roles mainly by agreement among sampled answers.