arXiv cs.CLSeptember 28, 2026
Likelihood Ranking doesn't Scale Like Prompting in LLMs
Excerpt
arXiv:2609.29390v2 Announce Type: replace Abstract: LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative stateme