← Back to all articles
arXiv cs.LGOctober 1, 2026

MedRECT: A Bilingual Medical Reasoning Benchmark for Error Correction in Clinical Texts

Excerpt

arXiv:2511.00421v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show promise in medical applications, but their ability to detect and correct errors in clinical texts remains under-evaluated, particularly beyond English. We introduce MedRECT, a bilingual benchmark for Japanese and English that formulates medical error handling as three subtasks: error detection, error sentence extraction, and error correction. MedRECT-ja contains 663 samples derived from the Japanese Medic