CVSD-Reg:用于鲁棒激光雷达配准的跨模态视觉语义先验蒸馏
CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration
AI总结:
本文提出CVSD-Reg框架,通过蒸馏视觉基础模型的语义先验提升激光雷达配准的鲁棒性,在多数据集上的成功率优于现有方法,且无需相机输入或事后ICP精修。
AI中文摘要:
基于学习的全局点云配准已取得显著进展,但现有方法依赖几何表示,对点密度、扫描模式、视角及传感器特性的变化敏感。本文提出CVSD-Reg,一种鲁棒的全局激光雷达配准框架,该框架从视觉基础模型中蒸馏视觉语义先验至激光雷达表示。第一阶段,Point Transformer V3学生模型通过对比蒸馏和球面流形对齐学习冻结的DINOv2教师模型,该过程保留教师嵌入空间的超球面几何;自监督InfoNCE一致性和软SE(3)不变性进一步促使生成视角鲁棒的描述符。第二阶段,蒸馏后的表示通过对应学习、感知密度的点丢失增强及端到端位姿优化适配配准任务。仅用单个检查点,CVSD-Reg可泛化至单传感器和零样本跨传感器场景,无需传感器特定适配,推理时完全无需相机。在KITTI、nuScenes和HeLiPR数据集上,CVSD-Reg的严格成功率(SR@0.5m/1°)分别达97.7%、99.0%和99.3%,其中稀疏16线Velodyne扫描的成功率为97.3%;其在不依赖相机输入或事后ICP精修的情况下,较最优几何配准方法最高提升44.0个百分点。
英文摘要:
Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations. In Stage 1, a Point Transformer V3 student learns from a frozen DINOv2 teacher through contrastive distillation and spherical-manifold alignment, which preserves the hyperspherical geometry of the teacher embedding space. Self-supervised InfoNCE consistency and soft $\mathrm{SE}(3)$ invariance further encourage viewpoint-robust descriptors. In Stage 2, the distilled representation is adapted to registration through correspondence learning, density-aware point-dropout augmentation, and end-to-end pose optimization. With a single checkpoint, CVSD-Reg generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference. On KITTI, nuScenes, and HeLiPR, CVSD-Reg achieves strict success rate (SR@0.5\,m/$1^\circ$) of 97.7$\%$, 99.0$\%$, and 99.3$\%$, respectively, including 97.3$\%$ on sparse 16-beam Velodyne scans. It outperforms state-of-the-art geometric registration methods by up to 44.0 percentage points without requiring camera inputs or post-hoc ICP refinement.