← Back to all articles
arXiv cs.LGOctober 2, 2026

Does Scaling Reinforcement Learning Really Require More Training?

Excerpt

arXiv:2610.01133v1 Announce Type: new Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-fr