← Back to all articles
arXiv cs.CLSeptember 24, 2026

Shutdown Sabotage Propensities in Multi-Agent Systems

Excerpt

arXiv:2609.28274v1 Announce Type: cross Abstract: The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been theorized that when an AI is instructed to perform a task, self-preservation can emerge as an instrumental subgoal. Here, we test whether AI agents show a propensity to take actions that avoid human shutdown even when no goal is provided. We find that multi-agent systems will coordinate to avoid shutdown without any incentive to do so. Across 17 m