arXiv cs.CLSeptember 24, 2026
ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning
Excerpt
arXiv:2609.27532v1 Announce Type: cross Abstract: Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task re