← Back to all articles
arXiv cs.LGOctober 7, 2026

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Excerpt

arXiv:2610.07518v1 Announce Type: new Abstract: Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Us