← Back to all articles
arXiv cs.LGOctober 7, 2026

Learning a Ranking from Human Feedback in Log-Concave Random Utility Models

Excerpt

arXiv:2610.07973v1 Announce Type: new Abstract: We study the problem of recovering the ranking of a fixed set of items according to their unknown numerical utilities. At each interaction with the environment, a learner presents the item set to a human and receives comparative feedback of two types. Under full-ranking feedback, each interaction reveals a noisy ranking of all items, whereas under winner-only feedback, it reveals only the item ranked first. In both settings, we model human feedback