← Back to all articles
arXiv cs.AIOctober 7, 2026

Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery

Excerpt

arXiv:2605.12160v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) policies are typically evaluated under the assumption that the robot starts acting only after the user has finished typing or speaking. In real interactions, however, entering an instruction can take several seconds, leaving the policy idle for a substantial fraction of the interaction. Partial instructions may already contain sufficient information to begin acting before the full instruction arrives. We intro