arXiv cs.CLSeptember 18, 2026
An Analysis of Training-Free Self-Reported Confidence in Language Models
Excerpt
arXiv:2609.20541v1 Announce Type: new Abstract: Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\mathrm{True})$, and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is a surprisingly strong baseline: after auditing bench