arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23541cs.CVcs.RO

面向鲁棒多模态3D目标检测的视觉基础模型方法

Towards robust multimodal 3D object detection via visual foundation models

Ziying Song, Lin Liu, Hongyu Pan, Shaoqing Xu, Lei Yang, Mingzhe Guo, Caiyan Jia

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出RoboDistill框架,利用视觉基础模型SAM的领域微调、特征金字塔融合、深度引导小波注意力和知识蒸馏,提升多模态3D目标检测在27种分布外损坏下的鲁棒性。

中文摘要 AI 辅助

多模态3D目标检测是自动驾驶中鲁棒感知的基础,因为它整合了LiDAR和相机传感器的互补信息。然而,现有方法在由传感器噪声、恶劣天气和环境变化引起的分布外(OOD)损坏下,往往无法保持鲁棒性。为解决这一问题,我们提出了RoboDistill,一个鲁棒且可泛化的多模态3D目标检测框架,该框架利用视觉基础模型(VFMs),如Segment Anything Model(SAM)。首先,我们引入SAM-AD,一种领域特定的预训练策略,在自动驾驶图像上微调SAM,以提取具有丰富语义信息的特征表示。其次,我们设计了AD特征金字塔网络(AD-FPN),在多尺度上细化和上采样SAM特征,以便与LiDAR特征无缝融合。第三,我们开发了深度引导小波注意力(DGWA)模块,该模块抑制高频传感器噪声,同时保留关键上下文信息。最后,我们引入KD Fusion,其中预训练的SAM-AD作为教师,将高质量的视觉知识蒸馏到轻量级点云网络中,从而在噪声条件下提高鲁棒性。在27种具有挑战性的OOD损坏设置下进行的大量实验表明,RoboDistill相对于代表性的最先进方法,通常能提供更强或具有竞争力的检测性能和鲁棒性。这项工作弥合了VFM与3D目标检测之间的差距,并推动了真实世界自动驾驶应用中鲁棒多模态感知的发展。

英文摘要

Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and environmental changes. To address this problem, we propose RoboDistill, a robust and generalizable multimodal 3D object detection framework that leverages visual foundation models (VFMs), such as the Segment Anything Model (SAM). First, we introduce SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information. Second, we design the AD Feature Pyramid Network (AD-FPN) to refine and upsample SAM features at multiple scales for seamless fusion with LiDAR features. Third, we develop the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical contextual information. Finally, we introduce KD Fusion, in which the pretrained SAM-AD serves as a teacher that distills high-quality visual knowledge into a lightweight point-cloud network, thereby improving robustness under noisy conditions. Extensive experiments across 27 challenging OOD corruption settings show that RoboDistill generally delivers stronger or competitive detection performance and robustness relative to representative state-of-the-art methods. This work bridges the gap between VFMs and 3D object detection and advances robust multimodal perception for real-world autonomous-driving applications.

发表机构

  • Beijing Jiaotong University(北京交通大学)
  • Horizon Robotics(地平线机器人)
  • University of Macau(澳门大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑