arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DXPR:基于深度的视觉-激光雷达跨模态地点识别方法,利用视觉基础模型

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

Yungsoo Han, Youngseok Jang, Seungwon Roh, Jeongyeon Seo, H. Jin Kim

arXiv 2609.09005首次发表:更新:

发表机构

Seoul National University; KAIST(首尔大学; 韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DXPR框架,利用视觉基础模型将相机与激光雷达深度统一表示,实现跨模态地点识别,在KITTI和Boreas上性能优于现有基线。

AI 中文摘要

我们提出了DXPR,一种基于深度的跨模态地点识别(CMPR)框架,该框架利用视觉基础模型(VFM)将单目相机查询与激光雷达地图进行匹配,而无需特定模态的编码器。这使得机器人和自动驾驶车辆能够在预先构建的激光雷达地图中仅使用相机即可稳健地定位,即使在严重的季节、天气和光照变化下也是如此。关键思想是将相机图像和激光雷达扫描转换为统一的深度图像表示,从而使带有聚合头的单一VFM主干能够学习模态不变的全局描述符。为了使成对度量学习忠实于场景几何,我们引入了一种几何感知的重叠挖掘器:在相机和激光雷达深度进行跨模态尺度对齐后,我们在视图之间前向投影测量值以计算像素级重叠分数。该分数重新标记模糊对,并自适应地调节多相似性损失中的正边距,以避免在弱重叠视图上过拟合。在KITTI里程计和Boreas数据集上的大量实验证明了其在季节、天气和昼夜变化下的强大性能和鲁棒性。在KITTI上,DXPR在大多数序列上实现了近乎完美的Recall@1,并优于先前的CMPR基线。在Boreas上,DXPR在序列内性能上与强大的单模态基线(DINOv2-SALAD)相当,同时在更具挑战性的序列间设置中表现出明显改进。与RangeBEV相比,我们的方法在序列内和序列间评估中均持续表现更好,展示了在不同季节和光照变化下的鲁棒性。

英文摘要

We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR maps, even under severe seasonal, weather, and illumination changes. The key idea is to convert both camera images and LiDAR scans into a unified depth image representation so that a single VFM backbone with an aggregation head can learn modality-invariant global descriptors. To make pairwise metric learning faithful to scene geometry, we introduce a geometry-aware overlap miner: after cross-modal scale alignment of camera and LiDAR depth, we forward-warp measurements between views to compute a pixel-level overlap score. This score relabels ambiguous pairs and adaptively modulates the positive margin in a multi-similarity loss to avoid overfitting on weakly overlapping views. Extensive experiments on KITTI odometry and Boreas demonstrate strong performance and robustness across seasons, weather, and day/night. On KITTI, DXPR achieves near-perfect Recall@1 on most sequences and outperforms prior CMPR baselines. On Boreas, DXPR achieves intra-sequence performance on par with a strong single-modal baseline (DINOv2-SALAD), while showing clear improvements in the more challenging inter-sequence setting. Compared with RangeBEV, our method consistently performs better in both intra- and inter-sequence evaluations, demonstrating robustness under diverse seasonal and illumination changes.

Comments8 pages, 6 figures, and 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑