arXiv cs.LGOctober 7, 2026
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models
Excerpt
arXiv:2607.04332v2 Announce Type: replace Abstract: In this paper, we consider the setting where large language models (LLMs) are trained using reinforcement learning (RL) to simultaneously improve reasoning accuracy and verbalize their confidence. Our reward scheme uses two functions for rewarding confidence verbalized by the LLM: one for correct answers and the other for incorrect answers. If poorly designed, such a scheme may incentivize an LLM to answer incorrectly in order for its confidenc