← Back to all articles
arXiv cs.LGOctober 7, 2026

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces

Excerpt

arXiv:2606.06840v2 Announce Type: replace-cross Abstract: Reasoning-trained language models can perform, zero-shot, multi-label tasks that require selecting a small set of relevant labels from a universe of thousands to hundreds of thousands of candidates. We ask how they do it mechanistically, and whether the mechanism can be distilled. We make the question measurable by treating each decision as a token-level event scored by the model's own decision margin: the token that picks a coarse region