← Back to all articles
arXiv cs.LGOctober 2, 2026

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Excerpt

arXiv:2610.00388v1 Announce Type: new Abstract: Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Op