← Back to all articles
arXiv cs.AIAugust 18, 2026

SAPO: Step-Aligned Policy Optimization for Reasoning-Based Generative Recommendation

Excerpt

arXiv:2605.17648v2 Announce Type: replace Abstract: Generative recommendation treats next-item prediction as autoregressive item-identifier generation. Specifically, items are encoded as semantic identifiers (SIDs), which are short coarse-to-fine token sequences whose early tokens capture broad semantics and later tokens refine them. Recent work augments this paradigm with reasoning traces and optimizes them via reinforcement learning with verifiable rewards, typically outcome-reward algorithm w