← Back to all articles
arXiv cs.CLAugust 19, 2026

SOD: Step-wise On-policy Distillation for Small Language Model Agents

Excerpt

arXiv:2605.07725v3 Announce Type: replace Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model capacity. While reinforcement learning methods like group relative policy optimization provide only sparse outcome-level rewards. Recently, on-policy distillation (OPD) has gained popularity by supplying dense token-level supervision from a teacher on student-generated trajectories. However, our e