arXiv cs.CLSeptember 11, 2026
"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated
Excerpt
arXiv:2508.05830v3 Announce Type: replace Abstract: Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" evaluations and demonstrate an applied case of criterion contamination. N = 110 participants completed both structured diagnostic depression interviews (Mirror condition) and life history interviews ("Non-Mirror" condition