← Back to all articles
arXiv cs.LGOctober 2, 2026

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Excerpt

arXiv:2610.02179v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher si