arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SV-WAM:一种用于端到端自动驾驶的高效环视世界-动作模型

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

Jinyang Wang, Shiwei Li, Junjian Wang, Zhiqiang Deng, Jianbin Gao, Yihang Zhao, Liu Liu, Yongjia Zhao, Jinlong Chen, Huirui Xu, Yifeng Pan, Kangwei Liu, Fan Ren, Ji Tao, Minghao Yang

arXiv 2609.03602首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; Chongqing Changan Technology Co., Ltd.; Civil Aviation University of China; Beihang University; Guilin University of Electronic Technology(中国科学院自动化研究所; 重庆长安科技有限公司; 中国民航大学; 北京航空航天大学; 桂林电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有驾驶世界模型空间覆盖受限且推理效率低的问题,提出SV-WAM模型,通过以动作为中心的因果掩码和可微分可行驶区域正则化器,在六摄像头输入下实现高效规划,在两个基准上达到最优规划性能并具备低延迟和零样本迁移能力。

AI 中文摘要

世界模型(WM)通过学习未来场景动态的预测表征,在端到端自动驾驶中展现出强大潜力。然而,推理时生成未来视频会带来巨大计算开销,导致近期许多驾驶WM采用单前置摄像头作为输入以实现高效部署。该设计在变道、汇流和转弯等安全关键操作中限制了空间覆盖范围。为解决这一局限,我们提出SV-WAM,一种环视世界-动作模型(WAM),其在保持高效推理的同时保留完整的六摄像头观测。SV-WAM将未来视频预测作为共享生成模型内动作学习的密集训练监督,而非推理时的输出。该设计的核心是一个以动作为中心的因果掩码,在联合动作-视频去噪过程中阻止动作token关注未来视频token。因此,部署时可舍弃视频分支,仅保留高效的动作规划。此外,我们引入一种可微分的可行驶区域合规性正则化器,惩罚车辆 footprint 角接近或越过可行驶边界,提升规划安全性与边界感知能力。在闭环NAVSIMv2基准和开环nuScenes基准上的大量实验表明,SV-WAM实现了最先进的规划性能,同时具备低推理延迟和有竞争力的零样本迁移能力。

英文摘要

World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as input for efficient deployment. This design restricts spatial coverage in safety-critical maneuvers such as lane changes, merges, and turns. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action-video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves state-of-the-art planning performance with low inference latency and competitive zero-shot transfer capability.

Comments23 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑