← Back to all articles
arXiv cs.AIOctober 7, 2026

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Excerpt

arXiv:2607.19691v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since intermediate reasoning must be externalized as natural-language tokens. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despit