← Back to all articles
arXiv cs.LGOctober 2, 2026

Fully Online Decentralized Learning in Stochastic Games with Unknown Independent Chains

Excerpt

arXiv:2610.01181v1 Announce Type: new Abstract: We consider stochastic games with independent controlled chains and unknown transition kernels, where players observe only their local states and realized payoffs. We develop a fully online, decentralized, and uncoordinated mirror-descent algorithm that operates in the dual space of occupancy measures for approximating stationary Nash equilibrium (NE) policies. The algorithm uses a single transition/reward sample at every primitive time step, relie