← Back to all articles
arXiv cs.LGOctober 2, 2026

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

Excerpt

arXiv:2610.00320v1 Announce Type: cross Abstract: Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and ben