arXiv cs.LGOctober 7, 2026
Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
Excerpt
arXiv:2610.07755v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM,