arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoGoal3D:基于3D感知融合与优化的协作式3D目标检测

CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement

Zhihao Yang, Zhiyu Xiang, Peng Xu, Tianyu Pu, Kai Wang, Eryun Liu, Dongping Zhang, Yong Ding

arXiv 2607.19036首次发表:更新:

发表机构

Zhejiang University; Zhejiang Provincial Key Laboratory of Multi-Modal Communication Networks and Intelligent Information Processing; China Jiliang University(浙江大学; 浙江省多模态通信网络与智能信息处理重点实验室; 中国计量大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对V2X协作目标检测中3D检测效果不佳的问题,提出CoGoal3D框架,通过两阶段管道提取和优化3D特征,设计融合模块与辅助任务,还提出数据增强策略,在多数据集上取得新的最优性能。

AI 中文摘要

V2X协作目标检测可通过聚合多个协作智能体的环境特征克服单车系统的局限性。然而,现有主流V2X感知方法主要关注2D BEV目标检测,在3D检测任务中,因忽略协作智能体高度和姿态差异导致的3D空间错位,结果较差。本文提出名为CoGoal3D的新型协作3D目标检测框架,在两阶段管道中逐步提取和优化3D特征。第一阶段设计多尺度3D感知全局融合模块减轻3D空间错位,第二阶段用3D点重建辅助任务优化结果提案。还提出有效的多智能体协作数据增强策略。在公共真实世界数据集上的大量实验表明,CoGoal3D取得了新的最优性能,在DAIR-V2X、V2V4Real和V2X-Real数据集上3D AP@0.7分别提高了10.86%、10.34%和10.18%。

英文摘要

V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from multiple collaborative agents. However, existing mainstream V2X perception methods mainly focus on 2D BEV object detection. When 3D detection task is concerned, inferior results are obtained because they ignore the 3D spatial misalignment caused by differing height and attitude among the collaborators. In this paper, we propose a novel collaborative 3D object detection framework called CoGoal3D, which extracts and refines the 3D feature gradually in a two-stage pipeline. In the first stage, a multiscale 3D-aware global fusion module is designed to mitigate the 3D spatial misalignment. The resulting proposals are then refined in the second stage with an auxiliary task of 3D point reconstruction. An effective multi-agent collaborative data augmentation strategy is further proposed to enrich the training data while minimizing information loss. Extensive experiments on public real-world datasets demonstrate that our CoGoal3D achieves new state-of-the-art performance, with 3D AP@0.7 improvements of 10.86%, 10.34%, and 10.18% on the DAIR-V2X, V2V4Real, and V2X-Real datasets, respectively. Code is available at https://github.com/Megalo-f/CoGoal3D.

CommentsAccepted to ECCV 2026. 17 pages, 8 figures, 6 tables. Code: https://github.com/Megalo-f/CoGoal3D

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑