← Back to all articles
arXiv cs.CLSeptember 18, 2026

What Does Privileged Information Add to On-Policy Self-Distillation?

Excerpt

arXiv:2609.20612v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view wit