arXiv cs.LGOctober 1, 2026
Spherical Interpolation for Backward-Compatible Multimodal Representations
Excerpt
arXiv:2609.39836v1 Announce Type: cross Abstract: Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive