← Back to all articles
arXiv cs.LGOctober 7, 2026

Does On-Policy Distillation for Safety Pose Backdoor Risks?

Excerpt

arXiv:2610.07654v1 Announce Type: new Abstract: On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored tea