arXiv cs.CLSeptember 24, 2026
Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models
Excerpt
arXiv:2609.27510v1 Announce Type: new Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following ability complicates task-based assessment. We introduce Uncheatable Eval, a dynamic benchmark that regularly collects newly published text to evaluate