← Back to all articles
arXiv cs.AIAugust 18, 2026

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

Excerpt

arXiv:2606.03988v3 Announce Type: replace Abstract: Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual represe