发表机构
Huazhong University of Science & Technology; Dongfeng Research & Development Institute(华中科技大学; 东风研发院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SimWAM是一种仅将视频生成用作训练信号的简单WAM,通过联合流匹配协同训练视频与动作专家,在NAVSIM上达91.5 PDMS且延迟更低,可零样本迁移至nuScenes,为高效自动驾驶提供可靠基线。
AI 中文摘要
世界动作模型(WAM)通过将视频动态先验迁移至动作预测来改进端到端自动驾驶,但现有方法在推理时需要代价高昂的未来帧生成。我们提出SimWAM,一种简单且有效的WAM,仅将视频生成用作训练信号。它通过联合流匹配协同训练预训练视频专家和轻量动作专家,采用独立注意力掩码使动作预测与未来帧无关,训练后可丢弃视频分支,仅保留能直接预测轨迹的独立规划器。由于两个专家无共享参数,仅通过统一注意力接口交互,视频骨干可替换,动作专家可独立扩展,无需修改学习目标或推理流程。我们进一步应用强化学习优化超越轨迹模仿的组合驾驶奖励。我们的SimWAM在NAVSIM上达到91.5 PDMS,以显著更低的延迟超越了基于WAM的最先进规划器,并零样本迁移至nuScenes。这些结果表明SimWAM是一个简单且可靠的基线,可轻易受益于视频生成的进展以实现高效自动驾驶。代码和模型权重可在此httpsURL获取。
英文摘要
In autonomous driving, World-Action Models (WAMs) have improved end-to-end planning by transferring video dynamics priors to action prediction, but many still couple planning with future-video generation at inference, incurring substantial computational overhead. We present SimWAM, a simple yet effective WAM that leverages future-video prediction solely as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without future-frame generation at inference. This design supports multiple pretrained video backbones and independent action-expert scaling within a shared attention interface, while preserving the joint learning objective. Moreover, we apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Experiments show that SimWAM achieves $91.9$ PDMS on NAVSIM with a favorable trade-off between accuracy and latency among world-model-based planners, while transferring zero-shot to nuScenes. It also achieves competitive planning accuracy on WOD-E2E and PhysicalAI-Autonomous-Vehicles. These results position SimWAM as a plain yet solid baseline for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.
CommentsThe code and model weights are available at https://github.com/H-EmbodVis/SimWAM/