arXiv cs.AIOctober 7, 2026
Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery
Excerpt
arXiv:2605.12160v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) policies are typically evaluated under the assumption that the robot starts acting only after the user has finished typing or speaking. In real interactions, however, entering an instruction can take several seconds, leaving the policy idle for a substantial fraction of the interaction. Partial instructions may already contain sufficient information to begin acting before the full instruction arrives. We intro