← Back to all articles
arXiv cs.LGOctober 7, 2026

What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization

Excerpt

arXiv:2610.07269v1 Announce Type: cross Abstract: Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLL