← Back to all articles
arXiv cs.LGOctober 1, 2026

Learning Process Rewards via Reasoning State Propagation

Excerpt

arXiv:2609.39220v1 Announce Type: cross Abstract: Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complement limited process supervision with scalable outcome supervision. However, existing PRMs often model reasoning prefixes independent