AI 中文总结
研究针对V2X协同感知算法发展受限问题,介绍基于CARLA模拟器的SimBEV2X数据生成工具及数据集,通过它建立基线并提出CoBEVFusion架构,实现上下文感知多智能体特征聚合,性能优越。
AI 中文摘要
车联网(V2X)通信中的协同感知可克服单个自动驾驶车辆的固有物理限制,但强大的V2X算法发展受限于缺乏大规模、多模态、多任务数据集,且真实世界多智能体数据收集和标注成本高昂。为此,我们引入基于CARLA模拟器构建的先进合成数据生成工具SimBEV2X,它能自动创建随机驾驶场景以收集多模态传感器数据及各类地面真值。我们还展示了SimBEV2X数据集,它是目前最大的V2X感知数据集。最后,我们使用CoopDet3D在该数据集上建立了强大基线,并提出CoBEVFusion架构,结合CoopDet3D与融合轴向注意力(FAX)实现上下文感知多智能体特征聚合,性能优越。
英文摘要
Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range. However, the development of robust V2X algorithms, particularly those relying on unified spatial representations like bird's-eye view (BEV) representation, is hampered by the lack of large-scale, multi-modal, multi-task datasets. Moreover, collecting and annotating a large set of synchronized, real-world multi-agent data is prohibitively expensive. This has resulted in a landscape where existing V2X datasets are notably limited in both size and scope. To overcome this, we introduce SimBEV2X, an advanced synthetic data generation tool built on the CARLA simulator. SimBEV2X automatically creates randomized driving scenarios to collect multi-modal sensor data alongside various types of ground truth including 3D bounding boxes with unique track IDs, HD map information, BEV segmentation maps, and semantic occupancy voxel grids from both vehicles and RSUs. We also present the SimBEV2X dataset, the largest V2X perception dataset to date. The dataset comprises 258 scenes, each involving up to 8 connected vehicles and up to 4 RSUs across a variety of road networks. The SimBEV2X dataset is an order of magnitude larger than existing V2X datasets and contains 102,200 frames, 588,520 lidar point clouds, more than 3 million images, over 27 million bounding boxes, and a comprehensive set of other annotations. Finally, we establish a strong baseline on the SimBEV2X dataset using CoopDet3D and propose CoBEVFusion, a novel architecture that combines CoopDet3D with fused axial attention (FAX) for context-aware multi-agent feature aggregation, resulting in superior performance. SimBEV2X, the SimBEV2X dataset, and CoBEVFusion are available at https://simbev2x.org and https://github.com/GoodarzMehr/SimBEV2X.
CommentsSubmitted to IEEE for review