← Back to all articles
arXiv cs.CLSeptember 22, 2026

Decomposing Error and Style in Automated Clinical Coding

Excerpt

arXiv:2609.24877v1 Announce Type: new Abstract: In automated clinical coding, where the label space spans tens of thousands of diagnosis and procedure codes, models are currently evaluated against a single gold annotation, treating any deviation as error. But we find when two teams code the same 110 ACI-Bench encounters, they agree on only 73% of codes (Jaccard similarity) for the same note; even after an independent clinical audit removes erroneous codes, agreement rises only to 77%. Is that ga