← Back to all articles
arXiv cs.AIOctober 2, 2026

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Excerpt

arXiv:2610.01458v1 Announce Type: new Abstract: Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a referenc