arXiv cs.AIOctober 7, 2026
ResOPD: Tail Residualization for Sparse On-Policy Distillation
Excerpt
arXiv:2610.04882v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse teacher interfaces, typically transmitting either the sampled-token score or a small Top-$k$ distribution. However, this sparse setting faces a fundamental dilemma: sampled-token estimators a