发表机构
School of Computer Science and Technology, Tongji University; Shanghai Research Institute for Intelligent Autonomous System, Tongji University; Shanghai Innovation Institute(同济大学计算机科学与技术学院; 同济大学上海智能自主系统研究院; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出GeoUniPR框架,通过构建几何一致性深度图像视图并引入SC-InfoNCE损失,在无需复杂对齐模块的情况下,实现了跨模态地点识别的最优性能与良好泛化能力。
AI 中文摘要
跨模态地点识别(CMPR)旨在识别异构传感模态(如视觉与激光雷达)下的同一地点。现有方法通常采用复杂的对齐模块、多阶段训练或预训练骨干网络的全微调来弥合模态差距。本研究从几何一致性视角重新审视CMPR,提出GeoUniPR,一个统一且简洁的几何一致性框架。GeoUniPR通过将激光雷达点云投影到相机视角以构建几何一致性深度图像视图(DIV),在表征层面缩小跨模态差异,建立RGB与激光雷达间的直接对应关系。我们进一步为DIV补充激光雷达原生线索,包括强度与表面法向量信息,生成多通道几何表征以提升结构一致性。基于该表征,GeoUniPR使用两个架构相同的模态特定ViT基编码器学习统一嵌入空间,通过参数高效适配进行训练,无需辅助对齐模块、多阶段训练或骨干网络全微调。此外,我们引入空间一致性InfoNCE(SC-InfoNCE),一种CMPR专用的对比损失函数,可抑制空间连续性下由距离引发的假阴性。在KITTI与KITTI-360数据集上的大量实验表明,GeoUniPR在同模态与跨模态地点识别中均达到了当前最优(SOTA)性能,且具备出色的跨数据集泛化能力。
英文摘要
Cross-modal place recognition (CMPR) aims to identify the same location across heterogeneous sensing modalities, such as vision and LiDAR. Existing methods commonly bridge the modality gap using complex alignment modules, multi-stage training, or full fine-tuning of pretrained backbones. In this work, we revisit CMPR from the perspective of geometric consistency and propose GeoUniPR, a unified and concise geometry-consistent framework. GeoUniPR reduces cross-modal discrepancy at the representation level by projecting LiDAR point clouds into the camera perspective to construct Geometry-Consistent depth image views (DIV), which establish direct RGB-LiDAR correspondence. We further augment DIV with native LiDAR cues, including intensity and surface-normal information, yielding a multi-channel geometric representation that improves structural consistency. Based on this representation, GeoUniPR learns a unified embedding space using two modality-specific ViT-based encoders with identical architectures, trained through parameter-efficient adaptation without auxiliary alignment modules, multi-stage training, or full backbone fine-tuning. In addition, we introduce Spatially-Consistent InfoNCE (SC-InfoNCE), a CMPR-specific contrastive objective that suppresses distance-induced false negatives under spatial continuity. Extensive experiments on KITTI and KITTI-360 demonstrate that GeoUniPR achieves state-of-the-art (SOTA) performance in both same-modal and cross-modal place recognition, with strong cross-dataset generalization.
Comments17 pages, 10 figures