arXiv cs.LGOctober 2, 2026
Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis
Excerpt
arXiv:2610.01896v1 Announce Type: new Abstract: Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-l