← Back to all articles
arXiv cs.LGOctober 1, 2026

Rising Multi-Armed Bandits with Known Horizons

Excerpt

arXiv:2602.10727v3 Announce Type: replace Abstract: Rising Multi-Armed Bandits (RMABs) model sequential decision problems where each arm's expected reward improves with repeated pulls. In such problems, the value of investing in an arm depends on how much time remains, making knowledge of the horizon useful side information, yet its benefit remains underexplored. We investigate this benefit through CURE-UCB, a horizon-aware algorithm that estimates each arm's cumulative reward over the remaining