Reddit r/MachineLearningSeptember 19, 2026
Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D]
Excerpt
Since OpenAI retired SWE-bench Verified in February (every frontier model tested could reproduce reference fixes for some tasks; underspecified tests rewarded knowing the intended fix), I've been trying to write down precisely what a decontamination report can and can't establish. The claim: training-side decontamination has a hard floor. A report is a claim by the party whose score depends on it, over a corpus nobody else can inspect, using matching that misses paraphrase and synthetic derivati