← Back to all articles
arXiv cs.LGOctober 2, 2026

Backdoor Containment via Expert Quarantine and Shutdown in LLMs

Excerpt

arXiv:2610.00663v1 Announce Type: cross Abstract: Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weight