← Back to all articles
arXiv cs.AIOctober 7, 2026

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Excerpt

arXiv:2610.04272v1 Announce Type: cross Abstract: Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inference and environment interaction costs incurred by MOPD, we find that much of its reported accuracy gain is due to certain training