arXiv cs.CLSeptember 11, 2026
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
Excerpt
arXiv:2605.12227v3 Announce Type: replace Abstract: Existing approaches to post-train models for long-context tasks face complementary limitations: (i) supervised fine-tuning (SFT) provides stable supervision but suffers from exposure bias; (ii) reinforcement learning methods such as Group Relative Policy Optimization (GRPO) train on model-generated trajectories but struggle with long-horizon credit assignment and sparse rewards; and (iii) on-policy distillation (OPD) provides dense token-level