arXiv cs.LGAugust 17, 2026
Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
Excerpt
arXiv:2608.12973v2 Announce Type: replace-cross Abstract: In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scalin