arXiv cs.AIOctober 7, 2026
Cost-Effective Automated Judging of Natural-Language Mathematical Proofs
Excerpt
arXiv:2608.00004v2 Announce Type: replace-cross Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IMO-GradingBench, two of three cheap judges (GPT-OSS-120B, DeepSeek-V4-Flash) and their three-model consensus are sta