arXiv cs.LGOctober 1, 2026
Regret Analysis of Retry-Based Bandits
Excerpt
arXiv:2605.20854v3 Announce Type: replace Abstract: We provide the first regret analysis of ReMax in stochastic multi-armed bandits. Originally introduced for reinforcement learning, ReMax is motivated by the role of exploration when multiple attempts are allowed, and is closely related to retry-based objectives such as pass@$k$ and max@$k$, which value the best outcome across those attempts. Given a posterior over arm values, ReMax chooses a sampling distribution that maximizes the posterior ex