arXiv cs.LGOctober 2, 2026
Towards Optimal Policy Improvement
Excerpt
arXiv:2610.01566v1 Announce Type: new Abstract: Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single update under specified constraints. We show that optimal improvement restricted to a set of states is equivalent to solving an induced M