← Back to all articles
arXiv cs.AIAugust 18, 2026

TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents

Excerpt

arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approaches either rely on process evaluators, which incur annotation and inference costs, or derive step-level credit from successful trajectories. However, successful trajectories are extremely scarce during early-stage reinforcement learning, substantially we