arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06021cs.RO

结合视觉嵌入与前馈三维模型的拓扑自主车辆定位

Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models

Eulogio Quemada-Torres, Alberto Jaenal, Francisco-Angel Moreno, Javier Gonzalez-Jimenez

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出结合视觉嵌入与前馈三维模型的拓扑自主车辆定位框架,通过集成视觉地点识别与前馈三维几何模型的位姿估计,在三个基准上较现有方法有显著提升,且模块化特性使组件可互换。

中文摘要 AI 辅助

有效的视觉定位(VL)需要环境地图兼具紧凑性(保证高效可扩展性)、对视觉外观变化的鲁棒性以及度量精度。视觉地点识别(VPR)通过低维图像嵌入可满足前两项要求,但其度量精度较低,不如基于局部特征或神经表示的标准VL方法。这一局限可通过将VPR与前馈神经三维几何(FF3D)模型生成的精确局部轨迹估计相结合来克服。本文中,我们通过拓扑框架解决基于外观的序列定位问题,该框架在受控图像集中迭代结合概率VPR与FF3D度量位姿估计。我们的方法提出了一种自动离线建图工具,用于建模场景不同部分的拓扑位姿-外观交互,该地图后续由在线粒子滤波器使用,该滤波器根据里程计和地点置信度估计位姿,供FF3D推理使用,成功将神经度量估计融入基于概率外观的定位。我们在三个已知基准上对该框架进行了广泛评估,结果表明其相较于现有基于外观的方法有显著提升。我们方法的模块化特性使描述符提取器和FF3D模型可互换,进一步的聚焦分析显示,序列置信度可缓解感知混淆下的严重失效问题。

英文摘要

Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition (VPR) is able to successfully meet the first two requirements, but its low metric accuracy makes it less suitable than standard VL approaches based on local features or neural representations. This limitation can be overcome by integrating VPR with the accurate local trajectory estimates produced by feed-forward neural 3D geometry (FF3D) models. In this paper, we address sequential appearance-based localization through a topometric framework that iteratively combines probabilistic VPR with FF3D metric pose estimation in controlled image sets. Our approach proposes an automatic offline mapping tool that models the topometric pose-appearance interaction in the different parts of the scene. This map is later employed by an online particle filter that estimates the pose from odometry and belief over places for FF3D inference, successfully incorporating neural metric estimation into probabilistic appearance-based localization. We extensively evaluate the framework on three known benchmarks, demonstrating substantial improvements over existing appearance-based methods. The modularity of our approach allows the descriptor extractor and FF3D model to remain interchangeable, and a focused analysis further shows that sequential belief can mitigate severe failures under perceptual aliasing.

发表机构

  • University of Malaga(马拉加大学)
  • University of Zaragoza(萨拉戈萨大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑