arXiv cs.LGOctober 2, 2026
ROGUE: Evaluating Corrigibility Failures in Frontier Computer-Use Agents
Excerpt
arXiv:2606.00341v2 Announce Type: replace Abstract: As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become paramount. Although much work has focused on agent safety in the presence of an adversary, we study corrigibility: whether agents remain amenable to human correction, interruption, or shutdown while pursuing benign tasks. We introduce ROGUE, a bench