arXiv cs.LGOctober 2, 2026
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
Excerpt
arXiv:2610.00348v1 Announce Type: cross Abstract: Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been ident