← Back to all articles
arXiv cs.LGOctober 2, 2026

CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

Excerpt

arXiv:2610.00636v1 Announce Type: cross Abstract: Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achi