arXiv cs.CLAugust 19, 2026
The Authenticity Gap in Human Evaluation
Excerpt
arXiv:2205.11930v3 Announce Type: replace Abstract: Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores. However, little consideration has been given as to whether this approach faithfully captures human preferences. Analyzing this standard protocol through the lens of utility theory in economics, we identify the implicit assumptions it makes about annotators.