AI 中文总结
该研究提出协同联合感知与预测(Co-P&P)框架,对比不同融合策略性能,实现端到端原型,发现检测级融合性能更优,协同可提升预测准确率且神经压缩能大幅降低带宽。
AI 中文摘要
联网自动驾驶车辆(Connected Autonomous Vehicles, CAVs)越来越多地利用车万物联网(Vehicle-to-Everything, V2X)通信来交换多源传感器信息,从而实现先进的协同感知(Collaborative Perception, CP)能力。除这些能力之外,本研究聚焦于协同联合感知与预测(Collaborative Joint Perception and Prediction, Co-P&P)这一范式,该范式将协同感知与运动预测相结合,以缓解两个长期存在的挑战:感知误差的累积和视觉遮挡。我们提出了协同联合感知与预测(Co-P&P)的概念框架,该框架可改善周围道路使用者的运动预测,从而提升复杂动态交通环境中的态势感知能力。在前期研究的基础上,本扩展版本比较了不同融合策略的性能,并为感知与预测的模块化设计建立了基线性能。实验结果表明,与检测级或跟踪级融合相比,预测级融合会导致整体系统性能下降。我们还实现了一个极简的端到端Co-P&P原型,该原型通过RENO神经编解码器耦合协同点云共享,并通过FutureDet实现联合检测与预测,结果显示,协同作用可提升预测准确率,而神经压缩技术可在通信带宽降低约34倍的同时保留这一优势。
英文摘要
Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions. We present a conceptual framework for Collaborative Joint Perception and Prediction (Co-P&P) that improves motion prediction of surrounding road users, thereby enhancing situational awareness in complex and dynamic traffic environments. Building upon our preliminary study, this extended version compares the performance of different fusion strategies and establishes baseline performance for a modular design of perception and prediction. Experimental results show that prediction-level fusion leads to a decline in overall system performance compared to detection-level or tracking-level fusion. We further implement a minimal end-to-end Co-P&P prototype that couples collaborative point-cloud sharing via the RENO neural codec with joint detection-forecasting via FutureDet, showing that collaboration improves forecasting accuracy while neural compression preserves this benefit at roughly 34x lower communication bandwidth.
Comments28 pages, 4 figures, post-publication of conference paper, accepted at journal SN Computer Science