不仅仅是你在哪里:从跨视图定位中学习语义、结构和几何
More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Beijing Aerospace Control Center(北京航天飞行控制中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究如何在极端视角变化下通过跨视图定位让模型学习语义、结构和几何。提出CROSS框架,克服现有方法局限,经实验验证该框架在跨视图定位中性能最优,能有效学习多方面内容。
AI中文摘要:
在极端视角变化下实现一致的跨视图理解对空间智能至关重要,跨视图定位为此提供了途径。近期视觉基础模型使跨视图匹配更可行,但我们认为跨视图定位不应仅视为二维匹配或姿态估计。本文将其视为不止姿态估计的问题,研究其如何助力模型在极端视角变化下形成一致的跨视图理解,包括稳定语义、可靠结构和可转移几何。指出现有方法的三个关键局限,提出CROSS框架,基于三维对齐、结构感知匹配和假设排序。实验表明CROSS在跨视图定位中性能达最优,且有效学习了不同视角下的稳定语义、可靠结构和可转移几何。
英文摘要:
Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it requires a model to align ground-view imagery with geo-referenced satellite-view imagery despite drastic appearance changes to estimate camera poses. Recent visual foundation models have made this long-standing localization problem increasingly feasible by providing rich 2D representations for cross-view matching. However, we argue that cross-view localization should not be viewed merely as 2D matching or pose estimation. In this work, we revisit cross-view localization as more than pose estimation and investigate how it can help the model develop consistent cross-view understanding under extreme viewpoint changes, including stable semantics, reliable structure, and transferable geometry. We identify three key limitations of existing methods that prevent them from achieving this. They usually lack explicit 3D grounding, rely on strict point-wise matching that can weaken semantic consistency, and learn from an absolute objective that provides limited guidance for geometric reasoning. To address these limitations, we propose CROSS, a unified cross-view localization framework built upon 3D-grounded alignment, structure-aware matching, and hypothesis ranking. This formulation makes structure learning an intrinsic requirement, encourages semantic representations to remain stable, and enables the model to acquire transferable geometry. Extensive experiments on the KITTI and VIGOR datasets show that CROSS achieves state-of-the-art performance in cross-view localization. More importantly, CROSS effectively learns stable semantics, reliable structure, and transferable geometry across extremely different viewpoints.