← Back to all articles
arXiv cs.AIOctober 7, 2026

Transfer-Stratified On-Policy Distillation for RL-Improved Reasoning Teachers

Excerpt

arXiv:2610.05974v1 Announce Type: cross Abstract: Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher str