← Back to all articles
arXiv cs.CLSeptember 24, 2026

Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks

Excerpt

arXiv:2609.27900v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in multi-agent systems where a principal agent decomposes tasks and delegates them to subordinate agents that may invoke external tools. Safety alignment, however, is still evaluated almost exclusively under a single-agent threat model, treating safety as a property of the individual LLM. We show that this assumption breaks down: \emph{individual safety alignment fails to transfer to multi-agen