发表机构
Tianjin University; Peking University(天津大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对自动驾驶中激光雷达与相机融合检测依赖2D预训练骨干网络的问题,提出DeGuNet,通过稀疏感知机制有效对齐图像与激光雷达深度,实验证明其能消除架构冗余,提升效率并提高mAP,建立参数高效多模态3D感知新范式。
AI 中文摘要
在自动驾驶感知中,激光雷达和相机模态融合已成为3D目标检测的主导范式。但当前多模态框架严重依赖在2D语义任务上预训练的大量视觉骨干网络,存在参数冗余和结构错位问题。为此提出DeGuNet,一种专为深度引导表示学习设计的超紧凑即插即用图像骨干网络。通过稀疏感知特征提取机制,有效对齐多视图图像与非结构化激光雷达深度,防止无效区域污染。在nuScenes数据集上的实验表明其广泛适用性和高效性,集成到现有基线中可消除架构冗余,减少GPU内存消耗达66.5%,推理速度加快1.16倍,mAP增益达6.20,建立了参数高效多模态3D感知的新范式。
英文摘要
In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current multi-modal frameworks heavily rely on massive visual backbones pretrained on 2D semantic tasks. This reliance introduces substantial parameter redundancy and a structural misalignment, as 2D priors are ill-equipped to handle the extreme sparsity of LiDAR projections required for Bird's-Eye-View geometry. To address this, we present DeGuNet, an ultra-compact and plug-and-play image backbone explicitly designed for depth-guided representation learning. By incorporating sparsity-aware feature extraction mechanisms, DeGuNet effectively aligns multi-view images with unstructured LiDAR depth while strictly preventing invalid-region contamination. Extensive experiments on the nuScenes dataset demonstrate DeGuNet's broad plug-and-play applicability and superior efficiency. When integrated into established baselines, it fundamentally eliminates architectural redundancy, reducing GPU memory consumption by up to 66.5% and achieving a 1.16x inference speedup. Concurrently, DeGuNet delivers up to a 6.20 absolute mAP gain, establishing a new paradigm for parameter-efficient multi-modal 3D perception.
CommentsAccepted to ECCV 2026