arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12854cs.ROcs.AIcs.CV

BrainWAM:面向自动驾驶的语义先验与预测动力学的动作空间协调框架

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

  • Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所)
  • Li Auto Inc.(理想汽车)

机构由 AI 辅助整理,请以论文原文为准。

Bing Zhan, Shuyao Shang, Shuo Lu, Yuan Xu, Zhao Wang, Yida Wang, Xueyang Zhang, Kun Zhan, Jiahao Gu

中文总结 AI 辅助

该研究针对自动驾驶中语义与预测动力学的规划需求,提出BrainWAM框架,通过结构化动作空间协调及异步整流流推理,在NAVSIM数据集上实现最优性能,优于仅VLA或仅WAM方法。

中文摘要 AI 辅助

自动驾驶需要在语义约束和预测动力学的双重条件下进行规划。然而,现有的端到端驾驶方法通常仅强调该需求的某一方面:视觉-语言-动作(VLA)模型利用视觉语言模型(VLM)先验进行语义推理,而世界动作模型(WAM)则通过生成式世界建模提供未来感知的预测。这自然催生了一种能够同时利用语义先验和预测动力学的统一规划器。但我们发现,通过联合令牌级注意力进行的朴素组合会出现注意力分配不匹配问题,即语义捷径会主导共享注意力空间,抑制预测动力学。受神经科学中复杂行为源于功能特化系统间协调的证据启发,我们提出BrainWAM,这是一种结构化动作空间协调框架,将语义推理和预测世界建模转化为两个专门的面向动作的通路,并在紧凑的动作表示层面对齐它们。我们进一步引入异步整流流推理策略,该策略具有解耦的视频和动作去噪功能,可缩短推理延迟,同时保留与规划相关的预测上下文。BrainWAM在NAVSIM v1(89.5 PDMS)和NAVSIM v2(89.6 EPDMS)上均达到了最先进的性能,始终优于仅VLA或仅WAM方法,凸显BrainWAM是自动驾驶系统的实用且有前景的方向。

英文摘要

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find that a naive combination through joint token-level attention suffers from an attention-allocation mismatch, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics. Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. We further introduce an asynchronous rectified-flow inference strategy with decoupled video and action denoising, which shortens inference latency while preserving planning-relevant predictive context. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.

↑