arXiv cs.AIAugust 18, 2026
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
Excerpt
arXiv:2608.15286v1 Announce Type: cross Abstract: We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym. Across 2,128 evaluation runs spanning nine models in six families (four development, three pre-registered held-out, plus a frontier pass on two frontier-tier models that the pre-registration designates ex