← Back to all articles
arXiv cs.AIOctober 2, 2026

Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning

Excerpt

arXiv:2610.00296v1 Announce Type: cross Abstract: Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through controlled empirical evaluations across models and tasks. We distinguish two prediction targets: identifying questions a model is more likely to answ