arXiv cs.LGOctober 7, 2026
On-Policy Distillation with Negative-Policy Rollouts
Excerpt
arXiv:2610.07874v1 Announce Type: new Abstract: On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the s