← Back to all articles
arXiv cs.CLOctober 7, 2026

Visual Abstention in Unified Multimodal Models

Excerpt

arXiv:2610.07887v1 Announce Type: new Abstract: Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request p