arXiv cs.LGOctober 2, 2026
Q-Learning for Reachability in MEC-Free MDPs
Excerpt
arXiv:2610.01781v1 Announce Type: cross Abstract: Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-te