arXiv cs.CLOctober 7, 2026
Holdout Best-of-N: Unbiased Evaluation and Its Cost
Excerpt
arXiv:2610.08719v1 Announce Type: new Abstract: Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward. We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if $J<K$, for every pool size $M\ge N\ge2$. At $J=K-1$,