← Back to all articles
arXiv cs.CLSeptember 11, 2026

TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

Excerpt

arXiv:2609.11399v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevalence, this problem lacks dedicated benchmarks and systematic study. We analyze over 790,000 translation outputs from 12 LLMs across 22 language pairs (LPs) and identify 12 recurring