← Back to all articles
arXiv cs.LGOctober 1, 2026

STEPS: Selective On-Policy Self-Distillation for Reasoning

Excerpt

arXiv:2605.10194v2 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own reasoning trajectories. However, persistent all-token guidance can degrade training in our math RL setting. Such guidance may unnecessarily constrain non-critical tokens and reinforce biases induced by privileged information unavailable at inference. Motivated by these risks, we propose STEPS, a