← Back to all articles
arXiv cs.CLOctober 7, 2026

KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy

Excerpt

arXiv:2609.38480v2 Announce Type: replace Abstract: Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment.