← Back to all articles
arXiv cs.CLSeptember 22, 2026

When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits

Excerpt

arXiv:2609.24194v1 Announce Type: new Abstract: Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure. When presented with a terse correct solution and a commented buggy solution for the same MBPP problem, a public preference reward model selects the correct one no better than a coin flip (0.507). Subtracting the predictable surface component from such scores is increasingly common, b