← Back to all articles
arXiv cs.AIAugust 18, 2026

Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

Excerpt

arXiv:2608.15630v1 Announce Type: cross Abstract: The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the asses