arXiv cs.CLSeptember 10, 2026
Deep and shallow biases in language models
Excerpt
arXiv:2609.09901v1 Announce Type: new Abstract: Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend on a particular prompt wording. We introduce a bias depth score that measures both how strongly a model prefers its top answer under direct prompting and whether that answer survives scenario reframing. Across 4,442 opini