← Back to all articles
arXiv cs.LGOctober 1, 2026

Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning

Excerpt

arXiv:2605.15692v2 Announce Type: replace Abstract: We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that are observed at the start of each episode. Performance is measured by cumulative regret against the episode-wise optimal value, $\sum_{k=1}^K [V^{*,M^k} - V^{\pi^k,M^k}]$, where $M^k$ represents the action context in the $k$-th episode. We show that the MVP algorithm naturally extends to this framework and