arXiv cs.AIOctober 7, 2026
Same Pieces, Different Servers: A Tetris Benchmark for AI Agents as Served
Excerpt
arXiv:2603.02348v2 Announce Type: replace-cross Abstract: An agent meets a model as served: through an endpoint with a price card, a shared cache and other tenants, or on whatever hardware a self-hosted model runs. Benchmarks rank the weights. We introduce a Tetris benchmark that measures what agents get from models as served: every move is scored against an oracle, and every agent receives the same pieces. In five pre-specified experiments with nine open-weight models on one serverless provider