← Back to all articles
arXiv cs.LGOctober 1, 2026

Role-Adaptive Policy Optimization for Offline Reinforcement Learning

Excerpt

arXiv:2609.40149v1 Announce Type: new Abstract: Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning an