← Back to all articles
arXiv cs.AIOctober 7, 2026

ResOPD: Tail Residualization for Sparse On-Policy Distillation

Excerpt

arXiv:2610.04882v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse teacher interfaces, typically transmitting either the sampled-token score or a small Top-$k$ distribution. However, this sparse setting faces a fundamental dilemma: sampled-token estimators a