arXiv cs.AIOctober 7, 2026
The Implementation Lottery: Auditing Idea Reliability in Automated Research
Excerpt
arXiv:2607.26587v2 Announce Type: replace-cross Abstract: Automated research agents use program scores to judge ideas. We call variation in this evidence across implementations the implementation lottery. We introduce an Idea Reliability Audit that freezes mechanism specifications, samples independent programs, and compares selected code with fresh implementations of its mechanism. Across 3,048 assignments on 31 tabular classification tasks, all four primary aggregation tests have Holm-adjusted