← Back to all articles
arXiv cs.LGOctober 1, 2026

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning via a Zero-Mismatch Reference

Excerpt

arXiv:2605.14220v2 Announce Type: replace Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign different values to the same sequence under the same model weights, inducing Training-Inference Mismatch (TIM). TIM is difficult to inspect because it is entangled with off-policy drift and common stabilization mechanisms. In this wor