← Back to all articles
Reddit r/MachineLearningSeptember 9, 2026

What Sante's 83.83 on DiagnosisArena-MCQ actually measures [D]

Excerpt

https://preview.redd.it/xv2epabu6ioh1.png?width=1171&format=png&auto=webp&s=4e22c69855ec509bd63a038a24f765caeaafa59c Ant Ling reports 83.83 on DiagnosisArena-MCQ for Ling-3.0-flash-Sante, its new medical reasoning model. The suffix matters: the task provides case information, examinations and tests, then asks the model to choose from four diagnoses. That result tells us about selecting an answer when the candidate set and case evidence are supplied. It does not establish how the same model would