arXiv cs.LGOctober 1, 2026
Structure over Pixels: Learning Variable-Length Visual Programs
Excerpt
arXiv:2605.27696v3 Announce Type: replace-cross Abstract: Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates that control the length. We propose STROP, a discrete tokenizer that learns both a visual program and its image-dependent active length. A length head is trained with a four-phas