arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用高端街景图像进行视觉定位的精度潜力

Accuracy potential of visual localization exploiting high-end street-level imagery

Jonas Meyer, Stephan Nebiker, Pascal Theiler, Norbert Haala

arXiv 2607.24409首次发表:更新:

发表机构

FHNW(西北应用科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉定位精度潜力,引入结合多种技术的可扩展管道,利用FHNW Muttenz数据集评估,实验表明其平移和旋转精度可观,能补充GNSS定位,为3D地理空间数据采集等铺平道路。

AI 中文摘要

在自主导航、测量、机器人技术以及增强和混合现实等应用中,对相对于参考框架的准确可靠姿态信息的需求日益增长。视觉定位可作为全球导航卫星系统(GNSS)的补充定位方式,但其适用性和准确性往往受限。由于缺乏亚厘米级地面真值姿态的大规模户外公共数据集,视觉定位的精度潜力尚未得到系统研究。本文通过引入可扩展的视觉定位管道解决了这两个问题,该管道将精确地理参考的高分辨率街景图像直接用作场景表示,并结合先验引导的参考候选选择、实时局部运动结构重建和基于PnP的姿态估计。此外还介绍了FHNW Muttenz数据集,利用该数据集评估视觉定位的精度潜力,实验表明平移的中位姿态精度在1-5厘米范围内,旋转在0.05-0.1°范围内,有利条件下可低至1厘米和0.03°,结果表明视觉定位可补充测量级GNSS定位,为使用消费设备进行3D地理空间数据采集和全自动地理配准方法铺平道路。

英文摘要

Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, surveying, robotics, and augmented and mixed reality. Visual localization can serve as a complementary positioning modality to GNSS, whose applicability and accuracy are often limited. Yet, the accuracy potential of visual localization has not been systematically investigated against survey-grade demands. This is mainly due to the lack of publicly available, large-scale outdoor datasets with ground-truth poses in the sub-centimeter range. In this work, we address both gaps. We introduce a scalable visual localization pipeline that employs precisely georeferenced, high-resolution street-level imagery directly as the scene representation. It combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation. We further present the FHNW Muttenz dataset, a real-world dataset covering a contiguous 10 km street network mapped in two mobile mapping campaigns approximately 1.5 years apart. It consists of high-resolution reference imagery and query sequences acquired by four different cameras across five representative scenes. All images are precisely co-registered, yielding 6-DoF ground-truth poses in the sub-centimeter range. Using this dataset, we evaluate the accuracy potential of visual localization. Our experiments demonstrate median pose accuracies in the range of 1-5 cm for translation and 0.05-0.1° for rotation, reaching as low as 1 cm and 0.03° under favorable conditions. These results show that visual localization can complement survey-grade GNSS positioning, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches. The dataset is publicly available at: https://fhnw-muttenz-vl-dataset.github.io/.

Comments26 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑