arXiv cs.LGOctober 2, 2026
Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
Excerpt
arXiv:2610.00400v1 Announce Type: new Abstract: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive