← Back to all articles
arXiv cs.LGOctober 1, 2026

Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data

Excerpt

arXiv:2609.38880v1 Announce Type: cross Abstract: We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $\gamma$, write $H=(1-\gamma)^{-1}$, and let $\mu_{\min}$ and $t_{\operatorname{mix}}$ denote the minimum stationary probability and total-variation mixing time. We prove that last-iterate TD achieves sup-norm error at most $\varepsilon$ with high probability using$\widetilde O\left( \frac