arXiv cs.LGOctober 2, 2026
When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora
Excerpt
arXiv:2610.00202v1 Announce Type: cross Abstract: Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a policy could in principle learn the second instead of the first. We audit that possibility in GooseReason-0.7M. First we ask whether the asymmetry is visi