← Back to all articles
arXiv cs.AIOctober 7, 2026

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

Excerpt

arXiv:2610.06804v1 Announce Type: cross Abstract: A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters, but n