arXiv cs.LGOctober 7, 2026
Does On-Policy Distillation for Safety Pose Backdoor Risks?
Excerpt
arXiv:2610.07654v1 Announce Type: new Abstract: On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored tea