← Back to all articles
Reddit r/MachineLearningOctober 6, 2026

SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

Excerpt

We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the agent can't recover the fix from git history. https://preview.redd.it/0uopztlmpsth1.png?width=1200&format=png&auto=webp&s=59c0698ca0a76c17be41e41362324d6a262f1f25 Some findings: With one attempt per task