arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ViCo3D:利用视觉基础模型增强基于激光雷达的协作式3D目标检测

ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models

Haojie Ren, Songrui Luo, Lingfeng Wang, Yan Xia, Yao Li, Jing Li, Lu Zhang, Jiajun Deng, Yanyong Zhang

arXiv 2607.12959首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对V2X系统中基于激光雷达协作式3D感知的不足,提出ViCo3D框架。通过将点云投影为图像让VFM提取特征,引入融合模块及跨智能体融合策略,实现了先进的3D检测性能,提升了协作增益。

AI 中文摘要

车辆到万物(V2X)系统中基于激光雷达的协作式3D感知通常依赖于跨智能体融合鸟瞰图(BEV)特征。然而,当前的BEV表示通常由从头训练的激光雷达主干提取,以几何为主导且缺乏通用语义先验,限制了特征级协作的效果。视觉基础模型(VFM)在大规模图像数据上预训练,在学习2D任务的通用和信息丰富的视觉表示方面表现出强大能力,有潜力增强基于激光雷达的BEV表示以进行协作。但由于图像 - 点云模态差距大,将VFM应用于基于激光雷达的3D检测仍具挑战性。为此提出ViCo3D框架,从三方面进行适配:将点云投影到BEV平面作为三通道图像让DINOv2提取特征;在单智能体编码器中引入多尺度BEV融合模块;采用以自我为中心的跨智能体融合策略聚合信息。在DAIR - V2X和V2XSet上的实验表明ViCo3D实现了先进的3D检测性能,在DAIR - V2X上协作增益比先前方法高1.8倍,代码将公开。

英文摘要

LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents. However, current BEV representations, typically extracted by LiDAR backbones trained from scratch, are geometry-dominated and lack general semantic priors, inherently limiting the efficacy of feature-level collaboration. Meanwhile, vision foundation models (VFMs) pretrained on large-scale image data have demonstrated strong capability in learning general-purpose and informative visual representations for 2D tasks, and have the potential to enhance agent-wise LiDAR BEV representations for collaboration. Despite this potential, adapting VFMs to LiDAR-based 3D detection remains challenging due to the substantial image-point cloud modality gap. To bridge this gap, we propose ViCo3D, a collaborative 3D object detection framework powered by VFMs. Specifically, ViCo3D adapts VFMs to LiDAR-based collaborative perception from three aspects: First, ViCo3D projects point clouds onto the BEV plane as three-channel images, enabling DINOv2 to extract BEV-space visual features from LiDAR inputs. Besides, to effectively integrate these DINOv2-derived features with LiDAR geometric features, ViCo3D introduces a multi-scale BEV fusion module within the single-agent encoder. In addition, ViCo3D adopts an ego-centric cross-agent fusion strategy to aggregate complementary information from multiple agents. Experiments on DAIR-V2X and V2XSet demonstrate that ViCo3D achieves state-of-the-art 3D detection performance. Remarkably, it delivers up to 1.8x greater collaborative gains than prior methods on DAIR-V2X. The code will be made public available for future investigation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑