← Back to all articles
arXiv cs.LGOctober 7, 2026

E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation

Excerpt

arXiv:2610.05048v2 Announce Type: replace Abstract: On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is c