Video2Track:从真实世界交互视频到自动驾驶系统的可操控对抗性封闭赛道测试
Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems
浏览论文内容
中文总结 AI 辅助
本文提出Video2Track框架,通过两个耦合模块将真实驾驶视频交互场景转化为可操控对抗性封闭赛道测试,可复现实景交互场景并生成可控变体,为ADS验证提供可扩展方案。
中文摘要 AI 辅助
封闭赛道测试在自动驾驶系统(ADS)的验证与确认中发挥着基础性作用,尤其适用于安全关键场景,可在受控条件下实现可复现的评估。然而,现有多数方法仍依赖标准化协议或预定义轨迹,导致交互过于程式化,难以复现公共道路交通的自然复杂性。为解决这一局限,本文提出Video2Track框架,该框架将真实世界的交互驾驶场景从视频转化为可操控的对抗性封闭赛道测试。该框架包含两个紧密耦合的模块:第一个是场景语义映射模块,它利用视觉语言模型从驾驶视频中提取结构化语义,并通过检索增强生成将其映射到封闭赛道拓扑库,从而识别兼容的地图片段和交互锚点;第二个是动态交互测试模块,它基于已映射的拓扑和锚点,通过条件扩散模型生成多样化的多智能体轨迹,同时通过带参数化对抗目标的斯塔克尔伯格博弈调控交互强度。封闭赛道实验表明,所提框架可忠实地复现具有代表性的真实世界交互场景,并生成具有可控风险水平和交互风格的可执行场景变体,为实现真实且可操控的ADS验证提供了可扩展的方法。
英文摘要
Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical scenarios, by enabling reproducible evaluation under controlled conditions. However, most existing approaches still rely on standardized protocols or predefined trajectories, leading to overly scripted interactions and limited ability to reproduce the natural complexity of public-road traffic. To address this limitation, we propose Video2Track, a framework that transfers real-world interactive driving scenarios from videos into steerable adversarial closed-track testing. The framework consists of two tightly coupled modules. The first is a scenario semantic mapping module, which extracts structured semantics from driving videos using a vision-language model and grounds them onto a closed-track topology library via retrieval-augmented generation, thereby identifying compatible map segments and interaction anchors. The second is a dynamic interactive testing module, which conditions on the grounded topology and anchors to generate diverse multi-agent trajectories through a conditional diffusion model, while regulating interaction intensity via a Stackelberg game with a parameterized adversarial objective. Closed-track experiments demonstrate that the proposed framework can faithfully reproduce representative real-world interaction scenarios and generate executable scenario variants with controllable risk levels and interaction styles, providing a scalable approach for realistic and steerable ADS validation.