arXiv cs.AIOctober 7, 2026
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
Excerpt
arXiv:2610.06824v2 Announce Type: new Abstract: We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize e