arXiv cs.AIOctober 2, 2026
Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning
Excerpt
arXiv:2610.00296v1 Announce Type: cross Abstract: Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through controlled empirical evaluations across models and tasks. We distinguish two prediction targets: identifying questions a model is more likely to answ