← Back to all articles
arXiv cs.AIOctober 7, 2026

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

Excerpt

arXiv:2607.27420v2 Announce Type: replace Abstract: Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text