MapLightning:基于一维地图令牌的在线矢量化高清地图构建
MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens
浏览论文内容
中文总结 AI 辅助
MapLightning用一维地图令牌替代密集BEV网格,通过自注意力构建高清地图,减少中间令牌16.7倍,提升精度和速度,在nuScenes和Argoverse 2上达到最先进性能。
中文摘要 AI 辅助
在线矢量化高清地图构建对于实现安全自动驾驶至关重要,需要准确、实时的推理。先前的方法通常采用密集的鸟瞰视角(BEV)网格作为中间表示。我们提出MapLightning,用一组紧凑的一维可学习地图令牌替代密集的BEV网格。为了从图像特征构建地图令牌,我们选择自注意力而非标准交叉注意力,因为它能实现图像与地图令牌之间的联合交互和上下文聚合。我们基于Transformer的映射器将地图令牌与图像令牌拼接,应用全自注意力,丢弃图像令牌,并保留更新后的地图令牌用于解码。该设计具有三个优势。首先,我们的表示高效,使用更少的令牌,消耗更少的内存,运行更快。其次,轻量级设计允许地图解码器使用全交叉注意力而非可变形交叉注意力,以获得更好的全局上下文。第三,与基于BEV的方法不同,我们的网络不使用相机投影参数,因此对相机外参扰动具有鲁棒性。MapLightning比基于密集BEV的方法最多使用16.7倍更少的中间令牌,并在nuScenes和Argoverse 2上实现了最先进的准确性和效率。其轻量级变体在nuScenes上超过MapTRv2 +10.1 mAP,在Argoverse 2上超过+16.2 mAP,同时推理速度提升1.73倍(40+ FPS),内存减少53%。我们进一步展示了在不确定性感知地图构建和下游轨迹预测方面的改进。代码和模型将发布。
英文摘要
Online vectorized HD map construction is essential for scaling safe autonomous driving and requires accurate, real-time inference. Prior methods typically rely on dense bird's-eye-view (BEV) grids as the intermediate representation. We propose \textit{MapLightning}, which replaces the dense BEV grid with a compact set of 1D learnable map tokens. To construct map tokens from image features, we choose self-attention over vanilla cross-attention because it enables joint interactions and contextual aggregation among image and map tokens. Our transformer-based mapper concatenates map and image tokens, applies full self-attention, discards the image tokens, and retains the updated map tokens for decoding. This design offers three advantages. First, our representation is efficient, using fewer tokens, consuming less memory, and running faster. Second, the lightweight design allows the map decoder to use full rather than deformable cross-attention for better global context. Third, unlike BEV-based methods, our network does not use camera projection parameters, making it robust to camera-extrinsic perturbations. MapLightning uses up to 16.7$\times$ fewer intermediate tokens than dense BEV-based methods and achieves state-of-the-art accuracy and efficiency on nuScenes and Argoverse~2. Its lightweight variant surpasses MapTRv2 by +10.1 mAP on nuScenes and +16.2 mAP on Argoverse~2, while delivering 1.73$\times$ faster inference (40+ FPS) with 53\% less memory. We further show improvements on uncertainty-aware map construction and downstream trajectory prediction. Code and models will be released.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。