arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

S$^3$Geo:面向跨视角地理定位的结构-语义协同学习

S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization

Ziqian Mo, Hill Zhang, Haosheng Tan, Ling Li, Jiaheng Wei

arXiv 2610.11608首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Claremont McKenna College(香港科技大学(广州); 克莱蒙特麦肯纳学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对跨视角地理定位中结构与语义建模不足的问题,提出S$^3$Geo框架,通过DQP模块、OT对比学习与SKD策略协同建模,在两个数据集上实现优于现有方法的性能且推理复杂度未增加。

AI 中文摘要

跨视角地理定位(CVGL)旨在通过匹配不同视角(如无人机视角与卫星视角)拍摄的图像来估计地理位置。现有方法主要依赖视觉表示,但往往无法联合建模细粒度结构对应关系与语义先验,易在视觉相似但语义不同的区域间产生混淆,从而限制了大视角变化下的鲁棒性。为应对这些挑战,我们提出S$^3$Geo,一种用于跨视角匹配的结构-语义协同学习框架。具体而言,我们首先引入解耦查询池化(DQP)模块,从密集令牌中提取一组紧凑的区域感知特征,以显式建模局部结构模式;随后设计基于最优传输(OT)公式的查询级对比学习方案,在跨视角空间错位下建立软对应关系;此外,我们融入来自冻结CLIP教师的语义知识蒸馏(SKD)策略,以迁移语义先验与关系结构,从而提升对难负样本的区分能力。通过协同运作,语义先验提供鲁棒的上下文过滤,引导结构模块建立精确的空间对齐。在University-1652与SUES-200数据集上的实验表明,S$^3$Geo在不增加推理复杂度的情况下,始终优于现有最优方法,验证了联合建模结构与语义信息对CVGL的有效性。

英文摘要

Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views. Existing methods mainly rely on visual representations, but often fail to jointly model fine-grained structural correspondences and semantic priors, making them prone to confusion between visually similar but semantically different regions, and thus limiting robustness under large viewpoint variations. To address these challenges, we propose \textbf{S$^3$Geo}, a structure-semantic synergistic learning framework for cross-view matching. Specifically, we first introduce a Decoupled Query Pooling (DQP) module to extract a compact set of region-aware features from dense tokens, enabling explicit modeling of local structural patterns. We then design a query-level contrastive learning scheme with an optimal transport (OT)-based formulation to establish soft correspondences under cross-view spatial misalignment. Furthermore, we incorporate a Semantic Knowledge Distillation (SKD) strategy from a frozen CLIP teacher to transfer semantic priors and relational structures, thereby improving discrimination on hard negatives. By operating synergistically, the semantic priors provide robust contextual filtering, which guides the structural module to establish precise spatial alignments. Experiments on the University-1652 and SUES-200 datasets demonstrate that \textbf{S$^3$Geo} consistently outperforms state-of-the-art approaches without increasing inference complexity, validating the effectiveness of jointly modeling structural and semantic information for CVGL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑