arXiv cs.CLOctober 7, 2026
Language-model ratings of depression reflect the rater more than the patient
Excerpt
arXiv:2610.08501v1 Announce Type: new Abstract: Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two rand