← Back to all articles
arXiv cs.LGOctober 2, 2026

Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning

Excerpt

arXiv:2409.14557v5 Announce Type: replace-cross Abstract: We study a structured class of Markov Decision Processes, known as Exo-MDPs, in which the state space is partitioned into exogenous and endogenous components. Exogenous states evolve stochastically, independent of the agent's actions, while endogenous states evolve deterministically based on both state components and actions. Exo-MDPs capture many operations research settings, including inventory control, resource management, and ride-sha