← Back to all articles
arXiv cs.LGOctober 7, 2026

Privileged Context as Drift in On-Policy Self-Distillation

Excerpt

arXiv:2610.07842v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affect