arXiv cs.LGOctober 7, 2026
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
Excerpt
arXiv:2607.17524v2 Announce Type: replace-cross Abstract: We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document