← Back to all articles
arXiv cs.AIOctober 7, 2026

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning

Excerpt

arXiv:2605.09287v3 Announce Type: replace Abstract: Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensive tasks. However, existing methods encounter critical challenges in long-horizon credit assignment: (i) Reward Sparsity, where models receive only outcome feedback without step-level guidance to differentiate action quality; (ii) Isolated Credit, where credit is assigned to steps independently,