← Back to all articles
arXiv cs.CLSeptember 24, 2026

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

Excerpt

arXiv:2609.27532v1 Announce Type: cross Abstract: Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task re