arXiv cs.AIOctober 7, 2026
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Excerpt
arXiv:2605.28201v2 Announce Type: replace Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP context, causing harmful agentic behaviors such as unsafe actions or incorrect outputs. Existing studies typically focus on single-interaction attacks, where the agent observes adversarial content and immediately exhibits harmful be