← Back to all articles
arXiv cs.CLSeptember 22, 2026

Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics

Excerpt

arXiv:2609.23264v1 Announce Type: new Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted reviewing, where reviewers may use LLMs to improve clarity or presentation while preserving the underlying judgments. We propose a stat