← Back to all articles
arXiv cs.LGOctober 2, 2026

Towards Optimal Policy Improvement

Excerpt

arXiv:2610.01566v1 Announce Type: new Abstract: Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single update under specified constraints. We show that optimal improvement restricted to a set of states is equivalent to solving an induced M