← Back to all articles
arXiv cs.AIOctober 7, 2026

The Implementation Lottery: Auditing Idea Reliability in Automated Research

Excerpt

arXiv:2607.26587v2 Announce Type: replace-cross Abstract: Automated research agents use program scores to judge ideas. We call variation in this evidence across implementations the implementation lottery. We introduce an Idea Reliability Audit that freezes mechanism specifications, samples independent programs, and compares selected code with fresh implementations of its mechanism. Across 3,048 assignments on 31 tabular classification tasks, all four primary aggregation tests have Holm-adjusted