← Back to all articles
arXiv cs.AIOctober 2, 2026

A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

Excerpt

arXiv:2610.01938v1 Announce Type: cross Abstract: Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general