← Back to all articles
arXiv cs.AIOctober 7, 2026

MeSD: Multi-Evidence Self-Distillation for VideoLLM

Excerpt

arXiv:2610.06342v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A furthe