← Back to all articles
arXiv cs.AIOctober 2, 2026

The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

Excerpt

arXiv:2610.00054v1 Announce Type: cross Abstract: Reading an LLM judge's verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood-scoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds. The mechanism is that judges do not always lead with a verdict token, on 12% to 4