arXiv cs.LGAugust 17, 2026
Online Inference in Distributional Temporal-Difference Learning
Excerpt
arXiv:2608.14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges weakly to a centered Gaussian random element in Cram\'er space. We also prove that, conditionally on the observed trajectory, the