arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DPA-I2P:面向自动驾驶中图像与点云配准的深度引导投影对齐

DPA-I2P: Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration in Autonomous Driving

Wenxin Zhang, Hang Li, Zhiwei Xu, Qiankun Dong, Gang Wang, Tao Li

arXiv 2608.26589首次发表:更新:

发表机构

Nankai University; Haihe Lab of ITAI(南开大学; 海河人工智能与信息技术实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对自动驾驶中图像与点云配准的跨模态对应学习难题,提出DPA-I2P方法,通过RMDE、PVL与CQP提升配准性能,在KITTI、nuScenes数据集上较基线方法有显著提升,可迁移性更好。

AI 中文摘要

图像与点云配准旨在估计给定图像在三维场景点云中的相机位姿,这是自动驾驶及大规模室外定位的基础任务。近期的隐式对应学习方法通过在端到端框架中学习跨模态对齐提升了配准性能,实现了更精准的相机位姿估计。但由于图像与稀疏激光雷达点云之间存在固有的模态差异,可靠的跨模态对应学习仍具挑战性。为解决该问题,本文提出面向图像与点云配准的深度引导投影对齐方法(DPA-I2P)。与单纯的深度或特征拼接不同,光线条件度量深度编码(RMDE)与投影一致视觉提升(PVL)以结构化、几何感知的方式利用深度与视觉线索;此外,跨模态查询剪枝(CQP)在早期优化阶段抑制不可靠查询,以提升匹配稳定性。在KITTI与nuScenes数据集上的实验验证了所提方法的有效性:在KITTI上,DPA-I2P相比最强的隐式基线方法将RTE降低45.0%、RRE降低55.6%;在nuScenes上,DPA-I2P也较评估的基线方法提升了配准精度,表明其对不同驾驶场景具有更好的可迁移性。

英文摘要

Image-to-Point Cloud Registration aims to estimate the camera pose of a given image within a 3D scene point cloud, which is a fundamental task in autonomous driving and large-scale outdoor localization. Recent implicit correspondence learning methods have improved registration performance by learning cross-modal alignment in an end-to-end framework, leading to more accurate camera pose estimation. However, due to the inherent modality discrepancy between images and sparse LiDAR point clouds, reliable cross-modal correspondence learning remains challenging. To address this issue, we propose Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration (DPA-I2P). Unlike naive depth or feature concatenation, Ray-Conditioned Metric Depth Encoding (RMDE) and Projection-Consistent Vision Lifting (PVL) exploit depth and visual cues in a structured, geometry-aware manner. In addition, Cross-Modal Query Pruning (CQP) suppresses unreliable queries during early refinement to improve matching stability. Experiments on KITTI and nuScenes demonstrate the effectiveness of the proposed method. On KITTI, DPA-I2P reduces RTE and RRE by 45.0% and 55.6% over the strongest implicit baseline, respectively. On nuScenes, DPA-I2P also improves registration accuracy over the evaluated baselines, suggesting better transferability to different driving scenes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑