发表机构
Institute of Automation, Chinese Academy of Sciences; Chongqing Changan Technology Co., Ltd.; Civil Aviation University of China; Beihang University; Guilin University of Electronic Technology(中国科学院自动化研究所; 重庆长安科技有限公司; 中国民航大学; 北京航空航天大学; 桂林电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有驾驶世界模型空间覆盖受限且推理效率低的问题,提出SV-WAM模型,通过以动作为中心的因果掩码和可微分可行驶区域正则化器,在六摄像头输入下实现高效规划,在两个基准上达到最优规划性能并具备低延迟和零样本迁移能力。
AI 中文摘要
世界模型(WM)通过学习未来场景动态的预测表征,在端到端自动驾驶中展现出强大潜力。然而,推理时生成未来视频会带来巨大计算开销,导致近期许多驾驶WM采用单前置摄像头作为输入以实现高效部署。该设计在变道、汇流和转弯等安全关键操作中限制了空间覆盖范围。为解决这一局限,我们提出SV-WAM,一种环视世界-动作模型(WAM),其在保持高效推理的同时保留完整的六摄像头观测。SV-WAM将未来视频预测作为共享生成模型内动作学习的密集训练监督,而非推理时的输出。该设计的核心是一个以动作为中心的因果掩码,在联合动作-视频去噪过程中阻止动作token关注未来视频token。因此,部署时可舍弃视频分支,仅保留高效的动作规划。此外,我们引入一种可微分的可行驶区域合规性正则化器,惩罚车辆 footprint 角接近或越过可行驶边界,提升规划安全性与边界感知能力。在闭环NAVSIMv2基准和开环nuScenes基准上的大量实验表明,SV-WAM实现了最先进的规划性能,同时具备低推理延迟和有竞争力的零样本迁移能力。
英文摘要
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as input for efficient deployment. This design restricts spatial coverage in safety-critical maneuvers such as lane changes, merges, and turns. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action-video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves state-of-the-art planning performance with low inference latency and competitive zero-shot transfer capability.
Comments23 pages, 16 figures