arXiv cs.LGOctober 1, 2026
ReTaCo: Residual-Target Control for On-Policy Distillation
Excerpt
arXiv:2609.39275v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-$k$ probabilities to limit cost. Because EOPD renormalizes these probabilities, its tar