← Back to all articles
arXiv cs.LGOctober 7, 2026

CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks

Excerpt

arXiv:2610.07177v1 Announce Type: new Abstract: An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher