← Back to all articles
arXiv cs.AIOctober 7, 2026

Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression

Excerpt

arXiv:2610.04638v1 Announce Type: cross Abstract: Actor-critic methods reuse past experience to improve sample efficiency. However, historical data are typically regarded as off-policy samples for the current policy-improvement update. This work introduces Regularized Dual Averaging Actor Critic (RDA2C), which assigns a distinct role to replay. In regularized dual averaging, the subsequent policy is determined by accumulated policy-improvement feedback, so historical advantage estimates contribu