← Back to all articles
arXiv cs.AIOctober 7, 2026

Are We Measuring Scientific Intelligence? Rethinking the Evaluation of AI Scientists

Excerpt

arXiv:2610.04915v1 Announce Type: new Abstract: AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot