arXiv cs.CLOctober 7, 2026
The Missing Minimal Pair: Stereotype Evaluation in LLMs
Excerpt
arXiv:2610.08747v1 Announce Type: new Abstract: A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we p