← Back to all articles
Reddit r/MachineLearningSeptember 20, 2026

Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]

Excerpt

In February OpenAI stopped reporting SWE-bench Verified and recommended other labs stop too. Every frontier model they tested could reproduce the human-written reference fix, or verbatim details of the problem statement, for some tasks. Progress had slowed to six points in six months and it wasn't clear how much of the remaining score was capability at all. The lab that built the benchmark, and had every reason to keep it, is the one that retired it. The usual answer is a decontamination report: