arXiv cs.AIOctober 7, 2026
Readable Before Actionable: Causal Tracing of Indirect Prompt Injection
Excerpt
arXiv:2610.05295v1 Announce Type: new Abstract: Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choi