← Back to all articles
arXiv cs.LGOctober 7, 2026

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Excerpt

arXiv:2610.07723v1 Announce Type: cross Abstract: Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn promp