← Back to all articles
arXiv cs.CLSeptember 21, 2026

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

Excerpt

arXiv:2604.23478v3 Announce Type: replace Abstract: Large language models are widely used to judge the output of other language models, yet whether a judge returns the same verdict when the same request is worded differently remains largely unexamined. We study that question across four evaluation tasks and twenty-five judges from six providers. To support the analysis we release JudgeSense, a benchmark of 880 items from human-labelled corpora, each issued under two instructions that differ in w