← Back to all articles
arXiv cs.CLAugust 19, 2026

The Authenticity Gap in Human Evaluation

Excerpt

arXiv:2205.11930v3 Announce Type: replace Abstract: Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores. However, little consideration has been given as to whether this approach faithfully captures human preferences. Analyzing this standard protocol through the lens of utility theory in economics, we identify the implicit assumptions it makes about annotators.