← Back to all articles
arXiv cs.LGOctober 2, 2026

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

Excerpt

arXiv:2610.00348v1 Announce Type: cross Abstract: Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been ident