arXiv cs.LGOctober 2, 2026
Benchmarking Prompt Optimization of Large Language Models With Chess
Excerpt
arXiv:2610.00416v1 Announce Type: cross Abstract: Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution infrastructure. These challenges are amplified in automatic prompt optimization (APO), where evaluation is repeated throughout the search for better prompts. Studying APO therefore requires a benchmark that is cheap and d