← Back to all articles
arXiv cs.CLSeptember 21, 2026

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

Excerpt

arXiv:2609.21663v1 Announce Type: new Abstract: Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an English dataset for human-centered ASR evaluation. Using this dataset, we benchmark lexical metrics against several configurations o