← Back to all articles
arXiv cs.AIOctober 2, 2026

Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

Excerpt

arXiv:2610.01023v1 Announce Type: cross Abstract: Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejec