Reddit r/MachineLearningSeptember 20, 2026
Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]
Excerpt
In February OpenAI stopped reporting SWE-bench Verified and recommended other labs stop too. Every frontier model they tested could reproduce the human-written reference fix, or verbatim details of the problem statement, for some tasks. Progress had slowed to six points in six months and it wasn't clear how much of the remaining score was capability at all. The lab that built the benchmark, and had every reason to keep it, is the one that retired it. The usual answer is a decontamination report: