arXiv cs.AIOctober 7, 2026
Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
Excerpt
arXiv:2610.06804v1 Announce Type: cross Abstract: A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters, but n