AI 中文总结
针对现有驾驶世界模型基准缺乏车辆动力学验证的问题,提出基于CARLA-CarSim协同仿真平台、含10,080种配置的VehDyn基准,通过分层评估框架衡量轨迹、运动学和动力学一致性,发现现有模型动力学一致性不足,为物理一致的驾驶世界模型提供系统基础。
AI 中文摘要
视频世界模型正作为数据引擎、行动规划器和生成式模拟器应用于自动驾驶,但现有基准主要评估视觉保真度和粗略的物理合理性,对于生成的驾驶未来是否符合真实的车辆运动学和动力学,提供的证据有限。这一局限因缺乏车辆、道路、操作和速度条件独立控制,且真实车辆状态与视频同步记录的数据集而进一步加剧。我们提出VehDyn,一个面向车辆动力学的驾驶世界模型基准。VehDyn基于CARLA-CarSim协同仿真平台构建,该平台将逼真渲染与经过验证的多体动力学模型相结合,包含由五种车型、四种轮胎-路面摩擦系数、三种操作、四种目标速度、十四个场景和三种光照条件全因子设计产生的10,080种配置,每种配置均配有同步的位置、速度和姿态序列。基于该数据集,VehDyn引入了一个分层评估框架,用于衡量轨迹对齐、运动学一致性和动力学一致性,并对12种最先进的视频世界模型进行基准测试。我们进一步使用两种既定协议评估视频质量,并将其与VehDyn得分相关联。轨迹级指标接近饱和,十二个模型中有十个在真实值的20%以内,而没有一个模型在动力学一致性上达到真实值的92%,且视觉质量指标与车辆动力学保真度仅弱相关。DrivingWorld取得了最高的VehDyn得分,其次是Cosmos 3 Nano和LTX-Video 2.5,且VehDyn得分与人类判断高度一致。VehDyn为开发物理一致且视觉逼真的驾驶世界模型提供了系统性的基础。
英文摘要
Video world models are emerging as data engines, action planners, and generative simulators for autonomous driving, but existing benchmarks primarily assess visual fidelity and coarse physical plausibility, providing limited evidence on whether generated driving futures obey realistic vehicle kinematics and dynamics. This limitation is further compounded by the lack of datasets in which vehicle, road, maneuver, and speed conditions are independently controlled, and ground-truth vehicle states are recorded in synchrony with videos. We introduce VehDyn, a driving world model benchmark for vehicle dynamics. VehDyn is built on a CARLA-CarSim co-simulation platform where photorealistic rendering is coupled with a validated multi-body dynamics model, and it contains 10,080 configurations from a full factorial design over five vehicle types, four tire-road friction coefficients, three maneuvers, four target speeds, 14 scenes, and three illuminations, each paired with synchronized position, velocity, and attitude sequences. Built on this dataset, VehDyn introduces a hierarchical evaluation framework that measures trajectory alignment, kinematic consistency, and dynamic consistency, and benchmarks 12 state-of-the-art video world models. We further assess the video quality using two established protocols and correlate it with the VehDyn score. Trajectory-level metrics are nearly saturated, with ten of twelve models within 20\% of ground truth, while no model reaches 92\% of ground truth on dynamic consistency, and visual-quality metrics are only weakly correlated with vehicle-dynamics fidelity. DrivingWorld achieves the highest VehDyn score, followed by Cosmos 3 Nano and LTX-Video 2.5, and the VehDyn score agrees closely with human judgment. VehDyn provides a systematic foundation for developing driving world models that are physically consistent and visually realistic.
Comments48 pages, 27 figures, 19 tables