arXiv cs.CLSeptember 22, 2026
When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits
Excerpt
arXiv:2609.24194v1 Announce Type: new Abstract: Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure. When presented with a terse correct solution and a commented buggy solution for the same MBPP problem, a public preference reward model selects the correct one no better than a coin flip (0.507). Subtracting the predictable surface component from such scores is increasingly common, b