← Back to all articles
arXiv cs.LGOctober 1, 2026

MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies

Excerpt

arXiv:2609.39324v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared global action features may fail to establish timestep-specific correspondence between actions and local visual changes. To address