利用基于视觉的点云地图先验进行基于相机的3D目标检测与在线矢量化高清地图构建
Leveraging Vision-Based Point Cloud Map Priors for Camera-Based 3D Object Detection and Online Vectorized HD Mapping
浏览论文内容
中文总结 AI 辅助
提出用视觉构建的点云先验地图增强相机感知,通过Pi3X和DINOv3特征融合,在Argoverse 2上显著提升3D检测与矢量化地图精度,无需激光雷达。
中文摘要 AI 辅助
基于相机的3D目标检测和在线矢量化高清地图构建为自动驾驶提供了紧凑的场景表示,但两者都依赖于精确的度量几何,并且仍受深度模糊性的限制。在长期部署中,来自重复遍历的观测可以被累积成持久的点云先验,这些先验提供了超越当前观测的几何上下文。然而,现有的显式点云先验方法依赖于基于激光雷达的地图构建,因此需要昂贵的3D测距传感器。我们提出了一种框架,该框架使用Pi3X从先前的相机遍历中构建静态点云先验地图,并用DINOv3特征增强每个点。在运行时,使用全局定位检索局部先验补丁,用稀疏体素骨干网络进行编码,并在鸟瞰视图(BEV)中与提升的多视角相机特征融合。然后,任务特定的稀疏Transformer头从融合表示中预测3D目标和矢量化地图元素。在Argoverse 2上,基于视觉的先验将强基线从0.287提高到0.299的CDS,并将矢量化地图mAP从0.669提高到0.750。消融研究表明,语义DINOv3特征对矢量化地图构建尤为重要。这些结果表明,视觉构建的几何-语义先验为基于相机的感知提供了一种有效的长期场景记忆形式,无需激光雷达即可改善先验地图构建和在线推理这两个任务。
英文摘要
Camera-based 3D object detection and online vectorized HD mapping provide compact scene representations for autonomous driving, but both depend on accurate metric geometry and remain limited by depth ambiguity. Over long-term deployment, observations from repeated traversals can be accumulated into persistent point cloud priors that provide geometric context beyond the current observations. Existing explicit point cloud prior approaches, however, rely on LiDAR-based map construction and therefore require expensive 3D ranging sensors. We propose a framework that constructs a static point cloud prior map from previous camera traversals using Pi3X and augments each point with DINOv3 features. At runtime, a local prior patch is retrieved using global localization, encoded with a sparse voxel backbone, and fused in bird's-eye view (BEV) with lifted multi-view camera features. Task-specific sparse transformer heads then predict 3D objects and vectorized map elements from the fused representation. On Argoverse 2, the vision-based prior improves a strong baseline from 0.287 to 0.299 CDS and from 0.669 to 0.750 vectorized mapping mAP. Ablations show that semantic DINOv3 features are particularly important for vectorized mapping. These results demonstrate that vision-built geometric-semantic priors provide an effective form of long-term scene memory for camera-based perception, improving both tasks without LiDAR for prior-map construction or online inference.
发表机构
- University of Freiburg(弗莱堡大学)
机构由 AI 辅助整理,请以论文原文为准。