← Back to all articles
arXiv cs.AIOctober 7, 2026

How RL Reshapes LLM Reasoning: Transferability, Coverage, and Scaling Laws

Excerpt

arXiv:2610.04158v1 Announce Type: cross Abstract: Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve performance in other domains, yet gains in Pass@1 can coincide with lower Pass@$N$ than the base model. This raises a fundamental question: does RL expand an LLM's reasoning boundary, or merely reweight its existing reasoning space? We revisit these phenomena across Qwen and Gemma mode