AI 中文总结
该研究提出受人类驾驶启发的Driver2Map模型,可融合三种模态数据,通过两阶段对齐、姿态引导BEV融合等模块提升在线高清地图构建性能,在IoU和AP指标上优于现有方法。
AI 中文摘要
高清(HD)地图是自动驾驶系统的核心基础。在构建此类地图时,车载多视角相机图像、标准清晰度地图和卫星图像提供了关键信息。然而,由于这些数据源之间存在模态和视角差异,现有方法往往难以有效对齐与融合,导致在线高清地图构建仍具挑战性。为解决这些问题,我们提出Driver2Map——一种受人类驾驶启发的在线高清地图构建模型。与现有仅使用两种模态的高清地图构建模型不同,Driver2Map可同时利用三种模态。具体而言,我们提出“两阶段对齐”策略以减少不同模态间的空间错位;此外,我们引入“姿态引导的鸟瞰图(BEV)融合”模块,该模块利用相机姿态信息自适应加权多视角特征,从而在BEV生成过程中有效抑制跨视角特征重叠;同时,我们设计“用于地图优化的预训练先验”模块,通过学习地图结构先验优化初始预测,提升动态遮挡场景下的高清地图预测效果。大量实验表明,Driver2Map在交并比(IoU)和平均精度(AP)指标上均优于现有方法。
英文摘要
High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images provide crucial information. However, due to the modality and perspective differences among these data sources, existing methods often struggle to effectively align and fuse them, making online HD map construction still challenging. To address these issues, we propose Driver2Map, an online HD map construction model inspired by human drivers. Unlike existing HD map construction models that utilize only two modalities, our Driver2Map can simultaneously exploit three modalities. Specifically, we propose a "two-stage alignment" strategy to reduce spatial misalignment across different modalities. Additionally, we introduce "Pose-Guided BEV Fusion", a BEV (bird's-eye-view) generation module that leverages camera pose information to adaptively weight multi-view features, thereby effectively suppressing cross-view feature overlap during BEV generation. Also, we design a "Pretrained Prior for Map Refinement" module to refine the initial prediction by learning map structure priors, thus improving the HD map prediction under dynamic occlusions. Extensive experiments demonstrate that Driver2Map outperforms existing methods on both IoU and AP metrics.