arXiv cs.LGOctober 7, 2026
Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret
Excerpt
arXiv:2610.07737v1 Announce Type: cross Abstract: We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as $\mathrm{NR}_T = \mu^\star - (\prod_{t=1}^T \mathbb{E}\mu_{I_t})^{1/T}$, where $\mu_{I_t}$ is the mean reward of the recommended arm $I_t$ and $T$ is the horizon. Since it applies the geometric mean to per-round marginal expectations, it ignores the joint di