WM-Cov:交互式世界模型式自动驾驶仿真的测试充分性
WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
浏览论文内容
中文总结 AI 辅助
本文针对交互式世界模型式自动驾驶仿真测试的充分性问题,提出与提供方无关的 WM-Cov 评估层,经多组实验验证其可通过有效交互式证据收敛性更科学地评估测试效果。
中文摘要 AI 辅助
世界模型与生成式仿真器正成为自动驾驶领域新兴的交互式测试基础设施,因其可对 ego 规划器作出响应,生成反事实、罕见且关乎安全的 rollout。这将测试场景从固定的重放轨迹转变为交互式场景族,其实际演化取决于被测规划器。未解决的问题不仅是能否生成危险 rollout,还包括何种有效闭环证据足以支撑特定测试意图与停止决策。本文明确了交互式世界模型式测试充分性的定义,并提出 WM-Cov,一种与提供方无关的评估层,可将原始提供方输出转换为所需、已实现且有效的证据。WM-Cov 通过覆盖率增长、有效故障发现、故障模式多样性、真实性、伪影抑制、重复计数及有效证据精度来报告充分性。对已执行的 TeraSim/SUMO 事件、WM 类混合轨迹池及真实 DriveArena TrafficManager--WorldDreamer 矩阵的研究表明,看似危险的事件可能包含有效的 ADS 故障、重复项、部分实现及伪影。DriveArena 矩阵评估了 2 个规划器、2 个时域、6 个提示条件及 360 个 ego 路径请求;304 次尝试成为完全实现的证据,56 次仍为部分实现。不相交的 80 请求路径切片检查产生 74 次完全实现的尝试与 6 次部分实现的尝试。结果支持在预算约束下通过有效交互式证据的收敛性来评估世界模型式测试,而非仅通过原始生成的故障或提示覆盖率。
英文摘要
World models and generative simulators are emerging as interactive testing infrastructure for autonomous driving because they can react to the ego planner and produce counterfactual, rare, and safety-critical rollouts. This changes a test scenario from a fixed replayed trajectory into an interactive scenario family whose realized evolution depends on the planner under test. The unresolved question is therefore not only whether dangerous rollouts can be generated, but what valid closed-loop evidence is enough to support a specified testing intent and stopping decision. This paper formulates interactive world-model-style testing adequacy and introduces WM-Cov, a provider-agnostic evaluation layer that converts raw provider outputs into requested, realized, and valid evidence. WM-Cov reports adequacy through coverage growth, valid-failure discovery, failure-mode diversity, realism, artifact suppression, duplicate accounting, and valid-evidence precision. Studies on executed TeraSim/SUMO events, WM-like mixed trace pools, and a real DriveArena TrafficManager--WorldDreamer matrix show that dangerous-looking events can include valid ADS failures, duplicates, partial realizations, and artifacts. The DriveArena matrix evaluates two planners, two horizons, six prompt conditions, and 360 ego-route requests; 304 attempts become fully realized evidence and 56 remain partial. A disjoint 80-request route-slice check yields 74 fully realized and 6 partial attempts. The results support evaluating world-model-style testing by convergence of valid interactive evidence under budget, rather than by raw generated failures or prompt coverage alone.
发表机构
- School of Transportation Science and Engineering, Harbin Institute of Technology(哈尔滨工业大学交通科学与工程学院)
- Chongqing Research Institute of HIT(哈尔滨工业大学重庆研究院)
- Chongqing Changan Automobile Co., Ltd.(重庆长安汽车股份有限公司)
- University of Nis(尼什大学)
- University of Belgrade(贝尔格莱德大学)
机构由 AI 辅助整理,请以论文原文为准。