← Back to all articles
arXiv cs.AIOctober 7, 2026

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

Excerpt

arXiv:2510.00915v5 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) replaces costly human labeling with automated verifiers. To reduce verifier hacking, many RLVR systems binarize rewards to $\{0,1\}$, but imperfect verifiers inevitably introduce \emph{false negatives} (rejecting correct answers) and \emph{false positives} (accepting incorrect ones). We formalize verifier unreliability as a stochastic reward channel with asymmetric noise rates $\rho_0$