Sparse2comm:迈向鲁棒的协同3D目标检测
Sparse2comm: Towards Robust Cooperative 3D Object Detection
浏览论文内容
中文总结 AI 辅助
Sparse2comm提出带宽高效且鲁棒的协同3D目标检测框架,通过稀疏特征编码、延迟感知对齐和自校准融合,渐进恢复退化特征,在三个数据集上显著提升混合退化下的检测精度。
中文摘要 AI 辅助
协同感知通过共享车辆与路边基础设施之间的互补观测信息来提升自动驾驶中的3D目标检测性能。然而,实际部署受到有限带宽和不可靠协作的制约,其中数据包丢失、传输延迟和空间错位会共同劣化协同特征流。现有方法通常仅降低通信成本或补偿单一类型的退化,未能充分处理耦合干扰。为解决这一问题,我们提出Sparse2comm,一种带宽高效且鲁棒的协同3D目标检测框架,将不可靠协作视为对退化协同特征的渐进式恢复。稀疏特征编码首先将通信编码为协作智能体传输的随机掩码采样前景特征,自车从这些特征中重建稠密语义表示。这种稀疏到稠密的机制学习从稀疏观测中推断缺失的以目标为中心的内容,从而在同一表示中实现超低带宽通信和数据包丢失恢复。在语义恢复的特征上,延迟感知对齐预测运动流以补偿延迟消息,自校准融合在自适应跨智能体融合之前以自监督方式估计残余空间偏移。因此,Sparse2comm按有序流程恢复语义完整性、时间一致性和空间对齐。在DAIR-V2X、OpenV2V和V2V4Real上的大量实验表明,Sparse2comm在个体和混合真实世界退化下保持有竞争力的干净精度并持续提升鲁棒性。与选择性特征通信基线Where2comm相比,Sparse2comm在三个数据集上的混合设置AP@0.5/AP@0.7分别提升+20.15/+11.79、+12.66/+11.07和+15.36/+12.61。
英文摘要
Cooperative perception improves autonomous driving by sharing complementary observations among vehicles and roadside infrastructure for 3D object detection. However, practical deployment is constrained by limited bandwidth and unreliable cooperation, where packet loss, transmission delay, and spatial misalignment jointly degrade the cooperative feature stream. Existing methods often reduce communication cost or compensate for one degradation type, leaving coupled disturbances insufficiently addressed. To address this problem, we propose Sparse2comm, a bandwidth-efficient and robust cooperative 3D object detection framework that treats unreliable cooperation as progressive restoration over degraded cooperative features. Sparse Feature Encoding first encodes communication as randomly mask-sampled foreground features transmitted by collaborating agents, from which the ego vehicle reconstructs dense semantic representations. This sparse-to-dense mechanism learns to infer missing object-centric content from sparse observations, enabling ultra-low-bandwidth communication and packet-loss recovery within the same representation. On the semantically restored features, Latency-Aware Alignment predicts motion flow to compensate delayed messages, and Self-Calibrating Fusion estimates residual spatial offsets in a self-supervised manner before adaptive cross-agent fusion. Sparse2comm therefore restores semantic completeness, temporal consistency, and spatial alignment in an ordered pipeline. Extensive experiments on DAIR-V2X, OpenV2V, and V2V4Real show that Sparse2comm maintains competitive clean accuracy and consistently improves robustness under individual and mixed real-world degradations. Compared with the selective feature communication baseline Where2comm, Sparse2comm improves mixed-setting AP@0.5/AP@0.7 by +20.15/+11.79, +12.66/+11.07, and +15.36/+12.61 on the three datasets, respectively.
发表机构
- Nanyang Technological University(南洋理工大学)
- Beihang University(北京航空航天大学)
- Beijing Institute of Technology(北京理工大学)
- Yanshan University(燕山大学)
- University of Macau(澳门大学)
- Korea Advanced Institute of Science and Technology(韩国科学技术院)
- University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。