← Back to all articles
arXiv cs.LGOctober 2, 2026

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

Excerpt

arXiv:2610.00317v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level reinforcement learning can address this limitation, but typically requires policy rollouts and closed-loop interaction, which are co