← Back to all articles
arXiv cs.LGAugust 17, 2026

Offline Deep Q* Estimation with Diffusion Models

Excerpt

arXiv:2608.14401v1 Announce Type: cross Abstract: In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. To address this issue, we propose a novel framework that decouples operator estimation from value function learning. In this approa