← Back to all articles
arXiv cs.AIOctober 7, 2026

Viva La Vida: Verification and Accumulation Failures in Multi-Agent Proof Search

Excerpt

arXiv:2610.04829v1 Announce Type: cross Abstract: When an agentic prover works on an open problem, there is no proof assistant to fall back on: its verifier and lemma library are ultimately language models judging model outputs. We instrumented such a system end to end and analyzed $51{,}754$ traced observations across three full runs ($186$ hours, \$$5{,}694$). We find three connected failure modes. First, the three-model verifier requires unanimity and treats parse or API failure as non-approv