← Back to all articles
arXiv cs.LGOctober 7, 2026

WASD: Wasserstein-based Knowledge Distillation for Large Language Models

Excerpt

arXiv:2610.07706v1 Announce Type: new Abstract: Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate