arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24094cs.RO

SIREN-Bench:行为驱动的应急车辆交互生成与评估

SIREN-Bench: Behavior-Driven Generation and Evaluation of Emergency-Vehicle Interactions

  • Rochester Institute of Technology(罗切斯特理工学院)
  • CVS Health(CVS健康公司)
  • City University of Hong Kong(香港城市大学)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

Yicheng Zhu, Tianmu Zhao, Haoxin Leng, Fan Zuo, Tao Li, Zilin Bian

AI总结:

该研究提出行为驱动的SIREN协同仿真平台,构建SIREN-Bench-v1基准,经3类任务评估揭示不同行为模式下的失效特性,为自动驾驶与交通安全研究提供可扩展平台。

AI中文摘要:

应急车辆(EMV)可通过周围民用车辆制动、变道或形成救援通道来重组周边交通。评估这些对安全至关重要的交互,需要对EMV特权和民用车辆响应进行行为级控制,同时配备一致的感知与真值数据。现有数据集和仿真基准无法直接提供这种组合。我们提出SIREN,一个行为驱动的SUMO-CARLA协同仿真平台,用于生成EMV与民用车辆的交互。SIREN将SUMO的网络级交通演化与行为逻辑,同CARLA的连续车辆控制及同步车载感知相结合;根据活动行为的不同,交互由SUMO、CARLA或二者共同控制。我们将该平台实例化为SIREN-Bench-v1,包含应急等级L1-L3下的7种参数化交互模板,以及3种行为族,配备同步传感器观测和模拟器原生标注。我们通过3项代表性任务展示该基准:3D目标检测、轨迹预测和视觉-语言风险理解。对9种轨迹预测器、4种基于LiDAR的检测器和5种视觉-语言模型的评估显示,存在行为依赖的失效模式:清交通交互对检测任务最难,特权路口通行对预测任务最难,平均而言没有任何学习型预测器优于恒速基准。视觉-语言模型在正常交通场景的表现远好于近事故和碰撞事件。这些结果证明了以行为为中心的基准测试的价值,并确立SIREN作为可扩展的数据生成与评估平台,适用于自动驾驶和交通安全性研究。

英文摘要:

Emergency vehicles (EMVs) can reorganize surrounding traffic as civilian vehicles brake, change lanes, or form rescue corridors in response to their passage. Evaluating these safety-critical interactions requires behavior-level control over both EMV privileges and civilian responses, together with consistent sensing and ground truth. Existing datasets and simulation benchmarks do not directly provide this combination. We present \textbf{SIREN}, a behavior-driven SUMO--CARLA co-simulation platform for generating EMV--civilian interactions. SIREN couples SUMO's network-level traffic evolution and behavior logic with CARLA's continuous vehicle control and synchronized onboard sensing; depending on the active behavior, the interaction is controlled by SUMO, CARLA, or jointly. We instantiate the platform as \textbf{SIREN-Bench-v1}, comprising seven parameterized interaction templates across emergency levels L1--L3 and three behavior families, with synchronized sensor observations and simulator-native annotations. We demonstrate the benchmark through three representative tasks: 3D object detection, trajectory prediction, and vision-language risk understanding. Evaluations of nine trajectory predictors, four LiDAR-based detectors, and five vision-language models reveal behavior-dependent failure modes. Traffic-clearance interactions are hardest for detection, privileged intersection traversal is hardest for prediction, and no learned predictor outperforms the constant-velocity reference on average. Vision-language models perform substantially better on normal traffic than on near-miss and collision events. These results demonstrate the value of behavior-centered benchmarking and establish SIREN as an extensible data-generation and evaluation platform for autonomous-driving and transportation safety research.

↑