← Back to all articles
arXiv cs.LGOctober 7, 2026

RELACE: retrospective likelihood-based action credit estimation for long-horizon language agents

Excerpt

arXiv:2610.07349v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain