arXiv cs.AIOctober 7, 2026
Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization
Excerpt
arXiv:2610.05230v1 Announce Type: cross Abstract: Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, and forces a tradeoff between long-horizon coverage and the local precision required for contact-rich manipulation. To address thes