发表机构
Sun Yat-sen University; The University of Western Australia(中山大学; 西澳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对跨视图地理定位的视角与外观差异问题,提出无几何扭曲的联合视图共识引导学习框架,在四个基准上实现最优性能。
AI 中文摘要
跨视图地理定位因街景图像与卫星图像间存在剧烈视角变化和巨大外观差异而颇具挑战性。现有方法常采用几何扭曲来呈现共视线索,但此类变换依赖严格的空间假设,且在视图依赖可见性下不可避免地引入严重视觉失真,导致监督信号含噪且对应关系脆弱。为解决该问题,我们提出一种全新的联合视图共识引导学习框架,完全规避显式几何扭曲。该框架不强制刚性空间对齐,而是在特征空间内动态挖掘并自适应强化语义共识。具体而言,训练期间的辅助联合视图通路支持直接跨视图交互,使每个视图能选择性聚合佐证证据至统一共识表示。为解决单视图与联合视图流间的特征异质性,我们引入全局模式探针作为语义字典,将不同模态投影至严格对齐的度量空间。在共识介导的对比目标引导下,训练期间单视图嵌入被明确拉向联合视图锚点,将此共识挖掘能力蒸馏至单视图编码器,以便推理时进行鲁棒检索。大量实验表明,我们的方法在四个标准基准上实现了最优性能,凸显了发现跨视图语义共识对可靠地理定位的重要性。
英文摘要
Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.