arXiv cs.AIOctober 7, 2026
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Excerpt
arXiv:2610.06401v1 Announce Type: cross Abstract: Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We