arXiv cs.AIOctober 7, 2026
Viva La Vida: Verification and Accumulation Failures in Multi-Agent Proof Search
Excerpt
arXiv:2610.04829v1 Announce Type: cross Abstract: When an agentic prover works on an open problem, there is no proof assistant to fall back on: its verifier and lemma library are ultimately language models judging model outputs. We instrumented such a system end to end and analyzed $51{,}754$ traced observations across three full runs ($186$ hours, \$$5{,}694$). We find three connected failure modes. First, the three-model verifier requires unanimity and treats parse or API failure as non-approv