← Back to all articles
arXiv cs.LGOctober 7, 2026

On-Policy Distillation with Negative-Policy Rollouts

Excerpt

arXiv:2610.07874v1 Announce Type: new Abstract: On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the s