arXiv cs.LGAugust 18, 2026
The Distributional View of Knowledge Distillation
Excerpt
arXiv:2608.15215v1 Announce Type: cross Abstract: Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler gradient is blind to which wrong token receives probability mass. We develop a distributional view in which the teacher is represented not by a single softened output but by a family of multi-temperature views - marginals of the annealing path of its logits - and the student is trained