arXiv cs.AIOctober 7, 2026
AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents
Excerpt
arXiv:2610.05140v1 Announce Type: new Abstract: As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain expertise, making it difficult to keep evaluation aligned with advances in agent capabilities. We address this challenge by investigating whether scientific-agent benchmarks can b