← Back to all articles
arXiv cs.LGOctober 2, 2026

SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning

Excerpt

arXiv:2610.00838v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relativ