← Back to all articles
arXiv cs.LGOctober 1, 2026

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Excerpt

arXiv:2609.39436v1 Announce Type: new Abstract: Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy distillation reaches high performance earlier in training and achieves both higher average performance during subsequent RLVR and hig