arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自然环境中基于几何条件的视觉地点识别

Geometry-Conditioned Visual Place Recognition in Natural Environments

Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya, David Hall, Kaushik Roy, Peyman Moghadam

arXiv 2609.27370首次发表:更新:

发表机构

CSIRO Robotics, CSIRO, Australia; Queensland University of Technology, Australia(澳大利亚联邦科学与工业研究组织机器人部门; 澳大利亚昆士兰科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自然环境中视觉地点识别因外观变化而失效的问题,提出深度感知蒸馏方法,利用几何基础模型推断的几何信息条件化视觉基础模型表示,在WildCross基准上显著提升Recall@1和Recall@5。

AI 中文摘要

自然环境中的视觉地点识别(VPR)因植被重复、显著地标稀少以及穿越过程中巨大的外观和视角变化而仍然具有挑战性。虽然同一地点的视觉观测可能变化显著,但其潜在的空间结构往往更为持久。我们通过深度感知蒸馏(DAD)利用这种互补的几何一致性,该方法在无需任何深度传感器的情况下,基于几何基础模型(GFM)推断出的几何信息,对预训练视觉基础模型(VFM)的令牌表示进行条件化处理。DAD并非将几何视为额外的输入模态,而是将图像对齐的深度投影到VFM令牌空间中,并通过通道级几何条件化选择性地调制视觉表示。一种两阶段的教师引导学习策略首先将几何条件化表示锚定到预训练的外观空间,然后针对地点判别进行细化。在WildCross基准上的评估显示,与匹配的外观基线相比,DAD将平均跨序列Recall@1从61.41%提升至66.37%,Recall@5从65.86%提升至72.49%,其中在反向穿越和长期外观变化下提升最大。这些结果表明,当视觉外观变得不可靠时,GFM派生的几何信息可以为VPR提供持久的结构先验。

英文摘要

Visual Place Recognition (VPR) in natural environments remains challenging due to repetitive vegetation, sparse distinctive landmarks, and substantial appearance and viewpoint variation across traversals. While visual observations of the same place can change considerably, their underlying spatial structure is often more persistent. We exploit this complementary geometric consistency through Depth-Aware Distillation (DAD), which conditions the token representations of a pretrained Vision Foundation Model (VFM) on geometry inferred by a Geometric Foundation Model (GFM), without any depth sensor. Rather than treating geometry as an additional input modality, DAD projects image-aligned depth into the VFM token space and selectively modulates visual representations through channel-wise geometric conditioning. A two-stage teacher-guided learning strategy first anchors the geometry-conditioned representation to the pretrained appearance space, before refining it for place discrimination. Evaluated on the WildCross benchmark, DAD improves average inter-sequence Recall@1 from 61.41% to 66.37% and Recall@5 from 65.86% to 72.49% over a matched appearance-only baseline, with the largest gains under reverse traversal and long-term appearance variation. These results show that GFM-derived geometry can provide a persistent structural prior for VPR when visual appearance becomes unreliable.

CommentsAccepted at the 28th International Conference on Digital Image Computing: Techniques and Applications (DICTA 2026). 8 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑