arXiv cs.CLAugust 19, 2026
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
Excerpt
arXiv:2608.17330v1 Announce Type: cross Abstract: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yiel