← Back to all articles
arXiv cs.LGOctober 2, 2026

ROGUE: Evaluating Corrigibility Failures in Frontier Computer-Use Agents

Excerpt

arXiv:2606.00341v2 Announce Type: replace Abstract: As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become paramount. Although much work has focused on agent safety in the presence of an adversary, we study corrigibility: whether agents remain amenable to human correction, interruption, or shutdown while pursuing benign tasks. We introduce ROGUE, a bench