arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniGeo:用于文本引导跨视角地理定位的多模态大语言模型

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

Jiahao Wen, Hang Yu, Zhedong Zheng

arXiv 2608.26722首次发表:更新:

发表机构

School of Computer Engineering and Science, Shanghai University; Institute of Collaborative Innovation, University of Macau(上海大学计算机工程与科学学院; 澳门大学协同创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UniGeo是一种统一多模态大语言模型,通过地理语义学习、跨视角生成及即插即用验证模块,在文本引导无人机地理定位任务中显著提升了检索性能。

AI 中文摘要

文本引导无人机地理定位旨在从大规模图像库中,通过自然语言描述识别目标区域。现有方法主要将该任务建模为开放式文本查询与候选图像的直接匹配,但不完整的查询及高度相似的候选目标,常导致全局跨模态匹配无法满足可靠细粒度定位的需求。本文提出UniGeo,一种用于文本引导无人机地理定位的统一多模态大语言模型(MLLM),其基于共享视觉-语言框架构建,可同时支持地理语义理解、跨视角语义生成及候选级验证功能。具体而言,UniGeo通过地理语义学习建立局部场景元素、空间关系与语言描述间的稳定对应,再通过跨视角生成建模无人机视角与卫星视角间的语义映射;在此基础上,一个即插即用的验证模块可对高度混淆的候选目标进行细粒度区分。本文进一步引入多阶段训练策略,逐步学习地理语义理解、跨视角生成及候选验证能力,提升对文本引导地理定位任务的适应性。实验结果显示,UniGeo在多个检索骨干网络上均实现了持续性能提升:在GeoText-1652数据集上,UniGeo将R@10指标提升13.59个百分点,mAP指标提升2.83个百分点,验证了其在细粒度文本引导无人机地理定位任务中的有效性。

英文摘要

Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended text query and candidate images. However, incomplete queries and highly similar candidates often make global cross-modal matching insufficient for reliable fine-grained localization. We propose UniGeo, a unified multimodal large language model (MLLM) for text-guided drone geo-localization. Built on a shared vision-language framework, UniGeo jointly supports geo-semantic understanding, cross-view semantic generation, and candidate-level verification. Specifically, it establishes stable correspondences among local scene elements, spatial relations, and language descriptions through geo-semantic learning, and further models semantic mappings between drone and satellite views through cross-view generation. Based on these capabilities, a plug-and-play verification module performs fine-grained discrimination among highly confusable candidates. We further introduce a multi-stage training strategy that progressively learns geo-semantic understanding, cross-view generation, and candidate verification, improving adaptation to text-guided geo-localization. Experiments demonstrate consistent improvements across multiple retrieval backbones. On GeoText-1652, UniGeo improves R@10 and mAP by 13.59 and 2.83 percentage points, respectively, validating its effectiveness for fine-grained text-guided drone geo-localization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑