arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07269cs.CVcs.LG

地方在语言中的留存:跨视角地理定位的零样本语言推理

What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization

Ayesh Abu Lehyeh, Jay Hwasung Jung, Safwan Wshah

首次发表
浏览论文内容

中文总结 AI 辅助

本文探究仅用语言进行跨视角地理定位的可行性,通过MLLM生成结构化描述并比较,发现描述在粗粒度先验下有效,但无法替代视觉细节,迈向可解释定位。

中文摘要 AI 辅助

跨视角地理定位通常被作为图像检索问题来解决,即通过联合训练的嵌入,将地面图像与卫星瓦片数据库进行匹配。这类模型准确,但需要大量成对监督,且无法展示支持匹配的证据。本文研究了一个不同的问题:仅通过语言能解决多少此任务?我们提示多模态大语言模型(MLLM)将每个地面全景图和每个卫星瓦片描述为结构化文本,并通过比较这些描述进行定位。所有组件均未训练。我们在来自美国四个城市的9,826对VIGOR数据上,于三种设置下进行评估。第一,描述忠实但不具区分性。它们在两个视角间高度一致,但按描述相似性对完整池排序几乎从未返回正确瓦片(Recall@1为0.39%)。第二,我们将池缩小至十个邻近瓦片,如同粗粒度先验所做。相同的描述现在变得有用:一个评估结构一致性的MLLM评判器使随机排序加倍,并匹配强词汇基线。它还能指出两个描述中哪些字段一致、哪些冲突,这是嵌入距离无法做到的,我们视之为迈向可解释定位的一步。第三,我们将评判器置于训练好的视觉检索器之上。在其排名错误的查询上,基于图像的重新排序有效,而基于我们描述的重新排序无效(Recall@1分别为23.5%和10.7%)。场景结构在转化为语言后得以保留,而区分邻近地点所需的精细外观细节则丢失。代码和提示词公开可用,见该https URL。

英文摘要

Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLLM) to describe each ground panorama and each satellite tile as structured text, and localize by comparing these descriptions. No component is trained. We evaluate on 9,826 VIGOR pairs from four U.S. cities, in three settings. First, the descriptions are faithful but not discriminative. They agree closely across the two views, yet ranking the full pool by description similarity almost never returns the correct tile (0.39% Recall@1). Second, we narrow the pool to ten neighboring tiles, as a coarse prior would do. The same descriptions now become useful: an MLLM judge that scores structural consistency doubles random ranking and matches a strong lexical baseline. It also states which fields of the two descriptions agree and which conflict, which an embedding distance cannot do, and which we see as a step toward interpretable localization. Third, we place the judge on a trained visual retriever. On the queries it ranks wrongly, reranking from images works, while reranking from our descriptions does not (23.5% against 10.7% Recall@1). Scene structure survives the conversion into language, while the fine appearance detail needed to separate nearby places does not. Code and prompts are publicly available at https://github.com/AyeshAbuLehyeh/GeoLingual.

发表机构

  • University of Vermont(佛蒙特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑