← Back to all articles
arXiv cs.LGOctober 2, 2026

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

Excerpt

arXiv:2610.00686v1 Announce Type: cross Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible token