← Back to all articles
arXiv cs.AIOctober 7, 2026

Learning to Read the Contextual Tokens in Diffusion Transformers

Excerpt

arXiv:2610.06844v1 Announce Type: cross Abstract: Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading this contextual space through natural-language interrogation. We train a lightweight bottleneck network that maps intermediate cont