arXiv cs.LGOctober 1, 2026
Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis
Excerpt
arXiv:2609.38851v1 Announce Type: cross Abstract: End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassis