arXiv cs.LGOctober 2, 2026
DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
Excerpt
arXiv:2610.00317v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are co