arXiv cs.LGOctober 2, 2026
Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning
Excerpt
arXiv:2409.14557v5 Announce Type: replace-cross Abstract: We study a structured class of Markov Decision Processes, known as Exo-MDPs, in which the state space is partitioned into exogenous and endogenous components. Exogenous states evolve stochastically, independent of the agent's actions, while endogenous states evolve deterministically based on both state components and actions. Exo-MDPs capture many operations research settings, including inventory control, resource management, and ride-sha