arXiv cs.AIAugust 18, 2026
Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning
Excerpt
arXiv:2608.16316v1 Announce Type: cross Abstract: Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates the transfer of the reasoning capabilities of large models to smaller, more efficient ones. On-Policy Distillation (OPD) offers a promising solution by matching output-token distributions along student-generated trajectories. However, video reasoning often depends on evi