arXiv cs.LGOctober 1, 2026
Trust the Critic More
Excerpt
arXiv:2609.39247v1 Announce Type: new Abstract: Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We