← Back to all articles
arXiv cs.CLSeptember 22, 2026

Efficient Reasoning Exploration via State-Conditioned Latent Steering with Progress Guidance

Excerpt

arXiv:2609.24066v1 Announce Type: new Abstract: Best-of-$N$ is a widely used inference strategy for complex reasoning, whose effectiveness depends on whether sampled candidates can cover diverse and high-quality reasoning paths. However, post-trained reasoning models often suffer from \emph{exploration collapse}, where independent rollouts repeatedly follow similar reasoning paths and limit the gains from increasing the rollout budget. Existing methods alleviate this issue by promoting broader e