arXiv cs.LGOctober 7, 2026
Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models
Excerpt
arXiv:2610.07247v1 Announce Type: new Abstract: Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios. On-policy self-distillation further reduces the reliance on external large teacher models while improving the reasoning ability of compact language models. However, existing methods typically either distill all token po