arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14185cs.CVcs.AI

基于双层路由与稀疏空间注意力的多视角BEV三维目标检测用于自动驾驶

Bi-Level Routing and Sparse Spatial Attention based Multi-View BEV 3D Object Detection for Autonomous Driving

Jing Zhang, Jiaqi Liu, Zibo Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对BEV多视角三维目标检测的计算复杂度和视角变换效率问题,提出Sparse-BEVNet,通过双层路由注意力、级联组注意力和稀疏空间交叉注意力,在nuScenes上mAP达45.2%,NDS达54.5%。

中文摘要 AI 辅助

基于鸟瞰图(BEV)的多视角三维目标检测面临计算复杂度高、多尺度特征提取困难以及密集的二维到BEV视角变换效率低下等挑战。为解决这些问题,本文提出了一种改进的BEV三维目标检测算法Sparse-BEVNet。首先,在图像特征提取网络中引入双层路由注意力(BRA)机制,以减轻骨干网络的计算负担。其次,在特征融合模块中采用级联组注意力(CGA),在不引入额外计算开销的情况下,增强不同层级特征之间的深度交互。此外,采用稀疏空间交叉注意力机制替代传统的密集视角投影流程。在公开的nuScenes数据集上的实验结果表明,所提方法实现了45.2%的平均精度(mAP)和54.5%的nuScenes检测分数(NDS),相较于基线模型分别提升了3.6%和2.8%。

英文摘要

Bird's Eye View (BEV)-based multi-view 3D object detection suffers from challenges of computational complexity, multi-scale feature extraction, and efficiency of dense 2D-to-BEV view transformation. To address these problems, this paper proposes an improved BEV 3D object detection algorithm Sparse-BEVNet. Firstly, a Bi-Level Routing Attention (BRA) mechanism is introduced into the image feature extraction network to reduce the computational burden of the backbone. Second, Cascaded Group Attention (CGA) is employed in the feature fusion module, which enhances deep interaction across features of different hierarchical levels without introducing additional computational overhead. Furthermore, a Sparse Spatial Cross-Attention mechanism is adopted to replace the conventional dense view projection pipeline. Experimental results on the public nuScenes dataset demonstrate that the proposed method achieves a mean Average Precision (mAP) of 45.2% and a nuScenes Detection Score (NDS) of 54.5%, corresponding to 3.6% and 2.8% improvements relative to the baseline model, respectively.

发表机构

  • Stuart Weitzman School of Design University of Pennsylvania(宾夕法尼亚大学斯图尔特·韦茨曼设计学院)
  • Courant Institute of Mathematical Science New York University(纽约大学库朗数学科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑