arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SafeGen:基于视觉语言模型的自动驾驶安全关键场景的目标条件视频扩散

SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

Jiangfan Liu, Zexuan Cui, Tianyuan Zhang, Zonglei Jing, Zonghao Ying, Yaoyuan Zhang, Jiakai Wang, Xiaoqi Jiang, Aishan Liu, Xianglong Liu

arXiv 2607.19701首次发表:更新:

发表机构

Beihang University; Zhongguancun Laboratory; Chery Automobile Co., Ltd.(北京航空航天大学; 中关村实验室; 奇瑞汽车股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对VLM在自动驾驶系统中的应用,提出SafeGen框架,将场景生成设为目标条件扩散过程,引入上下文接地最终状态推理和最终状态条件下的视频演化,实验表明该框架能提升评判得分,微调后可提高真实驾驶场景性能。

AI 中文摘要

视觉语言模型(VLM)越来越多地应用于自动驾驶系统,因此迫切需要在罕见但安全关键的场景下进行严格的安全评估。其中,与弱势道路使用者的交互是现实世界中故障的主要来源。然而,现有的安全关键场景生成方法主要依赖基于模拟器的管道,存在较大的模拟到现实差距,且往往无法捕捉现实、多样和不可预见的人车交互动态。我们提出了SafeGen,这是一种用于VLM自动驾驶中安全关键场景生成的目标条件扩散框架。我们的关键见解是将场景生成表述为目标条件扩散过程,其中预定义的灾难性最终状态作为强大的监督信号,引导生成自然演变为安全关键结果的时间连贯视频轨迹。在此基础上,我们引入了上下文接地最终状态推理,利用VLM分析良性驾驶上下文并推断人车交互中的潜在漏洞,产生诱导高风险场景的结构化最终状态规范。基于这些目标,我们进一步提出了最终状态条件下的视频演化,将语义威胁转化为物理上合理的视觉动态。具体来说,我们通过深度感知几何投影在场景中实例化高风险代理,然后进行边界条件扩散,以生成具有一致运动模式和时间连贯性的中间帧。在3个VLM自动驾驶系统上进行的广泛实验表明,与SoTA基线相比,SafeGen平均将使用VLM评判器评估VLM自动驾驶系统的理解和决策能力的评判总体得分提高了24.25%。此外,对VLM自动驾驶系统进行微调可使真实驾驶场景中的性能平均提高15.9%。

英文摘要

VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However, existing safety-critical scenario generation methods predominantly rely on simulator-based pipelines, which suffer from a substantial sim-to-real gap and often fail to capture realistic, diverse, and unforeseen human-vehicle interaction dynamics. We present SafeGen, a goal-conditioned diffusion framework for safety-critical scenario generation in VLMADs. Our key insight is to formulate scenario generation as a goal-conditioned diffusion process, where a predefined catastrophic end-state serves as a strong supervisory signal, guiding the generation of temporally coherent video trajectories that naturally evolve toward safety-critical outcomes. Building on this formulation, we introduce Context Grounded End State Reasoning, which leverages VLMs to analyze benign driving contexts and infer latent vulnerabilities in human-vehicle interactions, producing structured end-state specifications that induce high-risk scenarios. Conditioned on these targets, we further propose End State Conditioned Video Evolution, which grounds semantic threats into physically plausible visual dynamics. Specifically, we instantiate high-risk agents within the scene via depth-aware geometric projection, followed by boundary-conditioned diffusion to generate intermediate frames with consistent motion patterns and temporal coherence. Extensive experiments across 3 VLMADs demonstrate that SafeGen increases the Judge Overall Score, a metric using a VLM judge to evaluate VLMADs' understanding and decision-making, by 24.25% on average compared to SoTA baselines. Furthermore, fine-tuning a VLMAD improves performance in real-world driving scenes by an average of 15.9%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑