arXiv cs.LGAugust 18, 2026
Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits
Excerpt
arXiv:2510.22819v3 Announce Type: replace Abstract: The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learner's actual decisions and describes the evolution of the learning process over time. However, in multi-armed bandits, most existing algorithmic analyses mainly focus on the order of regret, while the last-iterate (simple regret) convergence rate remains less explored-