← Back to all articles
arXiv cs.AIOctober 7, 2026

On Semi-Markov Suboptimality in Hierarchical Reinforcement Learning

Excerpt

arXiv:2610.05338v1 Announce Type: cross Abstract: Hierarchical reinforcement learning uses temporally extended subtasks for exploration, yet committing to their execution can restrict both deployment and policy learning. We identify and separate the resulting execution and policy suboptimality. Task and execution trees distinguish reward objectives from policy choices and decision interruption. A Unified Value Function for HRL and a four-stage Generalized Hierarchical Bellman Equation then suppo