arXiv cs.AIOctober 7, 2026
MeSD: Multi-Evidence Self-Distillation for VideoLLM
Excerpt
arXiv:2610.06342v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A furthe