← Back to all articles
arXiv cs.CLSeptember 23, 2026

Calibration as a First-Class Criterion in LLM Evaluation

Excerpt

arXiv:2609.26489v1 Announce Type: new Abstract: Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy L