← Back to all articles
arXiv cs.LGOctober 1, 2026

Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification

Excerpt

arXiv:2609.38812v1 Announce Type: cross Abstract: Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments. Yet how trustworthy such self-verification is remains poorly understood. To investigate this question systematically, we introduce a diagnostic framework that identifies the first complete solution in each trajectory, determines whether it is objectively correct, and uses this ground truth to quan