arXiv cs.LGOctober 2, 2026
How Much Can Language Models Gain from Test-Time Computation?
Excerpt
arXiv:2610.01110v1 Announce Type: new Abstract: How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-PO