arXiv cs.LGOctober 1, 2026
Also Small Models Can Reasonably Self-Evaluate Their Confidence
Excerpt
arXiv:2609.39478v1 Announce Type: new Abstract: This study systematically evaluates self-evaluation-based uncertainty quantification across different language models of varying sizes on question-answering tasks spanning general to specialized knowledge domains. Using various self-evaluation methods where models judge their own predictions, we examine how model scale and domain specificity affect the quality of self-assessed confidence signals. Our results reveal that while accuracy predictably d