← Back to all articles
arXiv cs.CLAugust 19, 2026

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Excerpt

arXiv:2603.24472v4 Announce Type: replace Abstract: Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. However, in mathematical reasoning, we find that it can reduce response length while degrading performance. We trace this degradation to the suppression of epistemic verbalization - the model's expression of uncertainty during reasoning. Through controlled experiments varying conditioning context richness