发表机构
Zhongguancun Laboratory; Tianjin University; Beihang University; Nanyang Technological University; Tsinghua University; State Key Laboratory of Software Development Environment; University of Illinois Chicago(中关村实验室; 天津大学; 北京航空航天大学; 南洋理工大学; 清华大学; 软件开发环境国家重点实验室; 伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何生成逼真且可控的视觉交通场景,提出端到端条件扩散框架E2E-CDiff,基于前视图视觉观察联合去噪未来运动状态等,减轻规划控制不匹配,实验表明其在可控性与逼真性权衡上表现良好,还能引发挑战性交互,作为自我规划器有竞争力。
AI 中文摘要
生成既逼真又可控的闭环交通场景对于评估自动驾驶系统至关重要,尤其是在罕见的安全关键交互情况下。现有基于学习的方法往往难以平衡可控性和逼真性,要么对交通行为的细粒度控制有限,要么以行为合理性为代价生成可控场景。本文提出了E2E-CDiff,一种用于可控且逼真场景生成的端到端条件扩散框架。基于前视图视觉观察,E2E-CDiff联合对未来运动状态和与路线交互的背景车辆的可执行低级控制进行去噪。这种统一的状态-动作生成减轻了传统两阶段轨迹-然后-控制器管道中的规划-控制不匹配。可微引导进一步调节速度、确保可驾驶区域合规,并支持防撞或碰撞寻求行为,实现自然主义和安全关键场景生成。在Bench2Drive上的实验表明,与代表性的强化学习和模仿学习基线相比,E2E-CDiff实现了良好的可控性-逼真性权衡,而其碰撞引导变体在多个自动驾驶系统之间引发具有挑战性的交互。E2E-CDiff作为基于学习的自我规划器也具有竞争力,证明了端到端状态-动作扩散的通用性。
英文摘要
Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control over traffic behavior or controllable scenarios at the expense of behavioral plausibility. This paper presents E2E-CDiff, an end-to-end conditional diffusion framework for controllable and realistic scenario generation. Conditioned on front-view visual observations, E2E-CDiff jointly denoises future motion states and executable low-level controls for route-interacting background vehicles. This unified state-action generation mitigates the planning-control mismatch in conventional two-stage trajectory-then-controller pipelines. Differentiable guidance further regulates speed, enforces drivable-area compliance, and supports collision-avoidance or collision-seeking behaviors, enabling both naturalistic and safety-critical scenario generation. Experiments on Bench2Drive show that E2E-CDiff achieves a favorable controllability-realism trade-off compared with representative reinforcement- and imitation-learning baselines, while its collision-guided variant induces challenging interactions across multiple autonomous driving systems. E2E-CDiff also performs competitively as a learning-based ego planner, demonstrating the generality of end-to-end state-action diffusion.