← Back to all articles
arXiv cs.AIOctober 7, 2026

Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection

Excerpt

arXiv:2610.05163v1 Announce Type: cross Abstract: Long-horizon agents consume external content, invoke tools, and modify persistent state. Indirect prompt injection can exploit task-specific context, propagate across causally connected stages, and alter a consequential action while the workflow continues; we term this staged prompt injection. We build an automated, feedback-guided attack generation pipeline and apply it to Claude Code and Codex in their native runtimes. The confirmed attacks spa