arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SV2V-RSim:基于近真实数据的自选择性V2V协同感知综合基准

SV2V-RSim: A Comprehensive Benchmark for Self-Selective V2V Cooperative Perception with Near-Realistic Data

Yulu Wu, Chao Wei, Jujun Cheng, Zhangkai Ni, Haowen Wang, Dengyang Suo, Cong Chen, Xinyi Liu, Shangce Gao

arXiv 2609.32863首次发表:更新:

AI 中文总结

针对V2V协同感知数据集参与智能体少、选择策略静态及模拟与真实差距大的问题,提出近真实大规模数据集SV2V-RSim及自适应车辆选择模块SVA,优化性能与带宽权衡,验证高真实性和有效性。

AI 中文摘要

车对车(V2V)协同感知通过使车辆能够共享其直接视线之外的信息来增强自动驾驶能力。然而,现有的V2V数据集受到参与智能体数量少、协作者选择策略静态以及模拟环境与真实环境之间存在显著领域差距的限制。为克服这些挑战,我们引入了SV2V-RSim,这是一个大规模、多模态、近真实的模拟数据集,旨在提升智能体多样性和真实感。此外,我们提出了自适应车辆选择(SVA)模块,该模块优化协作者选择,以在感知性能与通信带宽约束之间取得平衡。我们的数据集使用基于虚幻引擎5的模拟器生成,该模拟器集成了高保真3D资产、多样化的环境和复杂的交通场景。自车指定范围内的所有车辆均配备传感器套件,从而实现动态和自适应的协作者选择。SV2V-RSim涵盖四张地图、四种天气条件、从日出到夜晚的六个时间段、203K帧激光雷达数据、402K帧RGB图像以及17个对象类别的788K个标注3D边界框,支持3D目标检测、分割和深度估计等一系列协同感知任务。对近期协同感知算法的基准测试表明,SVA实现了优越的性能-带宽权衡,而模拟到真实的实验和无参考图像质量评估验证了数据集的高真实性和实际有效性。我们的数据集和代码将公开发布。

英文摘要

Vehicle-to-Vehicle (V2V) cooperative perception enhances autonomous driving by enabling vehicles to share information beyond their direct line of sight. However, existing V2V datasets are limited by a small number of participating agents, static collaborator selection strategies, and a significant domain gap between simulated and real-world environments. To overcome these challenges, we introduce SV2V-RSim, a large-scale, multi-modal, near-realistic simulation dataset engineered to elevate agent diversity and realism. Additionally, we present the Select Vehicles Adaptively (SVA) module, which optimizes collaborator selection to balance perception performance against communication bandwidth constraints. Our dataset is generated using the Unreal Engine 5-based simulator that integrates high-fidelity 3D assets, diverse environments, and intricate traffic scenarios. All vehicles within a specified range of the ego vehicle are equipped with sensor suites, enabling dynamic and adaptive collaborator selection. SV2V-RSim encompasses four maps, four weather conditions, six time periods from sunrise to night, 203K LiDAR frames, 402K RGB frames, and 788K annotated 3D bounding boxes across 17 object classes, supporting a range of cooperative perception tasks such as 3D object detection, segmentation, and depth estimation. Benchmarking on recent cooperative perception algorithms demonstrates that SVA achieves a superior performance-bandwidth trade-off, while sim-to-real experiments and No-Reference Image Quality Assessment validate the dataset's high realism and practical effectiveness. Our dataset and code will be publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑