← Back to all articles
arXiv cs.AIOctober 7, 2026

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

Excerpt

arXiv:2610.05622v1 Announce Type: cross Abstract: Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds a