← Back to all articles
arXiv cs.LGOctober 1, 2026

Sequential Bayesian Evaluation of Large Language Model Behavior

Excerpt

arXiv:2511.10661v2 Announce Type: replace-cross Abstract: It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assigned a binary or ordinal score and the aggregation of scores across prompts is then used as a summary evaluation. In this paper, we develop a Bayesian approach for quantifying the unc