arXiv cs.LGOctober 2, 2026
TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models
Excerpt
arXiv:2610.00899v1 Announce Type: cross Abstract: Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not ne