arXiv cs.LGOctober 2, 2026
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Excerpt
arXiv:2610.02117v1 Announce Type: cross Abstract: On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained percepti