arXiv cs.LGOctober 2, 2026
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
Excerpt
arXiv:2605.07579v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the