arXiv cs.LGOctober 2, 2026
Learning from the Near Future: Temporal Self-Distillation for RLVR
Excerpt
arXiv:2604.20733v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) is a core post-training recipe for reasoning models, yet pure on-policy learning can be inefficient when useful trajectories are difficult to discover or exploration narrows. Existing self-guided approaches largely reuse capability already available to the current or earlier learner. We instead ask whether learning can also make use of capabilities that emerge later in training: can a model