发表机构
Tsinghua University; Carnegie Mellon University(清华大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ControlPed框架,结合轨迹冲突合成与3D高斯泼溅,生成逼真可控的危险行人场景,评估显示端到端驾驶模型性能大幅下降。
AI 中文摘要
评估端到端自动驾驶在罕见、安全关键的车辆-行人交互场景中的表现,需要逼真的、传感器级别的场景。然而,基于轨迹的场景生成器无法合成原始视觉观测,而基于视频的方法则缺乏可控性。为弥合这一差距,我们提出了ControlPed,一种新颖框架,将轨迹级冲突合成与3D高斯泼溅(3DGS)相结合,以生成逼真、运动可控的安全关键场景。基于HazardPed数据集(该数据集源自10,352个交通视频,包含422条冲突轨迹、高清地图和857个带注释的3D人体运动),ControlPed首先生成冲突轨迹,通过文本条件运动扩散将其提升为3D人体运动序列,最后使用可动画化的3DGS化身渲染多视角传感器观测。在88个渲染的逼真场景中的安全评估显示,七个领先的端到端驾驶模型遭受严重性能下降,其平均HDScore从88.8骤降至47.4,暴露了在危险行人行为下的主要失败模式。数据集和测试基准将发布,以促进车辆-行人交互的安全评估。
英文摘要
Evaluating end-to-end autonomous driving under rare, safety-critical vehicle-pedestrian interactions requires photorealistic, sensor-level scenarios. However, trajectory-based scenario generators cannot synthesize raw visual observations, whereas video-based approaches lack controllability. To bridge this gap, we present ControlPed, a novel framework that combines trajectory-level conflict synthesis with 3D Gaussian Splatting (3DGS) to generate photorealistic, motion-controllable safety-critical scenarios. Built upon HazardPed, a dataset derived from 10,352 traffic videos comprising 422 conflict trajectories, HD maps, and 857 annotated 3D human motions, ControlPed first generates conflict trajectories, lifts them into 3D human motion sequences via text-conditioned motion diffusion, and finally renders multi-view sensor observations using animatable 3DGS avatars. Safety evaluation in 88 rendered photorealistic scenarios reveals that seven leading end-to-end driving models suffer a severe performance drop, with their mean HDScore plunging from 88.8 to 47.4, exposing major failure modes under dangerous pedestrian behaviors. The dataset and testing benchmarks will be released to facilitate safety assessment of vehicle-pedestrian interactions.
Comments9 pages, 7 figures, Website at https://controlped.netlify.app