← Back to all articles
arXiv cs.LGOctober 1, 2026

On the Complexity of Preference-Based Bandits

Excerpt

arXiv:2609.39351v1 Announce Type: cross Abstract: We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandi