← Back to all articles
arXiv cs.AIOctober 7, 2026

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

Excerpt

arXiv:2610.05140v1 Announce Type: new Abstract: As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain expertise, making it difficult to keep evaluation aligned with advances in agent capabilities. We address this challenge by investigating whether scientific-agent benchmarks can b