arXiv cs.LGAugust 18, 2026
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
Excerpt
arXiv:2506.03096v2 Announce Type: replace-cross Abstract: Contrastive language-image pre-training aligns features of text-image pairs in a common latent space via distinct encoders for each modality. While this approach achieves impressive performance in several zero-shot tasks, it cannot natively handle multimodal inputs, i.e., encoding image and text into a single feature vector. As a remedy, it is common practice to use additional modules to merge the features extracted by unimodal encoders.