arXiv cs.LGOctober 2, 2026
On Language Drift during RLVR Post-Training
Excerpt
arXiv:2610.02015v1 Announce Type: new Abstract: Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---