← Back to all articles
arXiv cs.LGOctober 2, 2026

TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models

Excerpt

arXiv:2610.00899v1 Announce Type: cross Abstract: Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not ne