arXiv cs.LGOctober 2, 2026
Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?
Excerpt
arXiv:2610.01399v1 Announce Type: new Abstract: Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving no