← Back to all articles
arXiv cs.CLOctober 7, 2026

The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

Excerpt

arXiv:2610.08026v1 Announce Type: new Abstract: In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them