← Back to all articles
arXiv cs.LGOctober 2, 2026

Benchmarking Prompt Optimization of Large Language Models With Chess

Excerpt

arXiv:2610.00416v1 Announce Type: cross Abstract: Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution infrastructure. These challenges are amplified in automatic prompt optimization (APO), where evaluation is repeated throughout the search for better prompts. Studying APO therefore requires a benchmark that is cheap and d