← Back to all articles
arXiv cs.LGOctober 2, 2026

When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

Excerpt

arXiv:2610.00202v1 Announce Type: cross Abstract: Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a policy could in principle learn the second instead of the first. We audit that possibility in GooseReason-0.7M. First we ask whether the asymmetry is visi