← Back to all articles
arXiv cs.CLSeptember 22, 2026

Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

Excerpt

arXiv:2609.22603v1 Announce Type: new Abstract: Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we intr