← Back to all articles
arXiv cs.AIOctober 7, 2026

Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning

Excerpt

arXiv:2610.05446v1 Announce Type: cross Abstract: Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal instead of the subgoal the high-level policy originally needed to update. We propose Hierarchical Time-aware Bootstrapping (HTB), which evaluates specified subgoals under the current low-level