arXiv cs.CLSeptember 14, 2026
Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models
Excerpt
arXiv:2609.12475v1 Announce Type: new Abstract: Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluation unnecessarily expensive. Although recent benchmark compression methods (BCMs) can mitigate this cost, many strong BCMs rely on large collections of per-sample evaluation results from numerous LLMs to identify representative samples. Building such collections is also expensive unless they are alrea