arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12557cs.CVcs.RO

DRS-VPT:利用视觉点变换器在扫描中直接重定位

DRS-VPT: Directly Relocalizing in a Scan with Vision Point Transformers

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Lanke Frank Tarimo Fu, Maurice Fallon

中文总结 AI 辅助

DRS-VPT提出一种前馈变换器架构,实现图像到扫描的直接配准,统一自动驾驶标定与室内重定位,达到最先进性能并支持零样本迁移。

中文摘要 AI 辅助

我们提出了DRS-VPT,一种用于基础图像到扫描配准的前馈变换器架构。给定查询图像和参考3D点云,该模型预测扫描姿态和点图,以及每个相机的姿态和点图,所有这些都在第一相机的坐标系中表示。它还预测了逐点和逐像素的从粗到细的特征金字塔,用于将扫描直接重投影对齐到第一张图像。这一公式统一了下游任务,如自动驾驶中的相机-激光雷达标定和室内相机到地图的重定位。单个DRS-VPT模型在自动驾驶中的图像到激光雷达配准上达到了最先进的性能,在无需训练地图特定权重的情况下实现了具有竞争力的室内重定位,并且对未见环境具有强大的零样本迁移能力。我们还定性地展示了该模型学习了复杂的扫描到图像投影属性,例如背面点的遮挡。

英文摘要

We present DRS-VPT, a feed-forward transformer architecture for foundational image-to-scan registration. Given query images and a reference 3D point cloud, the model predicts the scan pose and point map alongside the poses and point maps of each camera, all expressed in the first camera's frame. It additionally predicts a coarse-to- fine pyramid of per-point and per-pixel features for direct reprojective alignment of the scan to the first image. This formulation unifies downstream tasks such as camera-LiDAR calibration in autonomous driving and indoor camera-to-map relocalization. A single DRS-VPT model achieves state-of-the-art performance for image-to-LiDAR registration in autonomous driving, competitive indoor relocalization without training map-specific weights, and strong zero-shot transfer to unseen environments. We also show qualitatively that the model learns complex scan-to-image projection properties such as occlusion of back-facing points.

↑