arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DA-WAM:面向驾驶世界模型的决策对齐未来潜变量

DA-WAM: Decision-Aligned Future Latents for Driving World Models

Ruiguo Zhong, Benshan Ma, Xiaolong Chen, Lang Zhang, Mingyue Feng, Yaonong Wang, Pei Liu, Jun Ma

arXiv 2608.19085首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Leapmotor; The Hong Kong University of Science and Technology(香港科技大学(广州); 零跑汽车; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DA-WAM 是统一预测表示学习等模块的驾驶世界模型框架,通过动作条件未来潜变量实现决策对齐,在 NAVSIM 数据集上达到最优性能,验证了关键组件的有效性。

AI 中文摘要

预测在 ego 智能体的动作作用下场景如何演化是安全自动驾驶的基础,但驾驶世界模型在决策方面的全部潜力尚未被挖掘。核心挑战在于确保未来建模不仅具备预测性,还能为决策提供信息:预测的未来必须直接影响轨迹的选择。现有方法要么将未来表示学习与规划优化解耦,要么在多个轨迹候选间共享预测状态,从而削弱了本应指导选择的特定动作后果。为填补这一空白,我们提出 DA-WAM,这一框架在单一决策目标下统一了预测表示学习、动作条件未来建模和轨迹评分。DA-WAM 通过在线编码器和稳定的动量目标,在规划器优化全程保持预测监督,使未来表示能与驾驶任务协同演化。动作条件预测器为每个轨迹候选生成不同的未来潜状态,随后由未来潜变量条件化的因子化解码器进行评估。对于与专家匹配的轨迹,预测的未来潜变量会受到观测到的未来表示的监督,而安全关键的难例会在规划边界附近提供额外监督。在 NAVSIM-v1 和 NAVSIM-v2 上开展的大量实验表明,该方法达到了当前最优性能,而消融实验和诊断分析验证了关键组件的有效性。

英文摘要

Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of world models for decision-making remains unrealized. The critical challenge lies in ensuring that future modeling is not merely predictive, but decision-informative: the predicted future must directly shape which trajectory is selected. Existing approaches decouple future representation learning from planning optimization, or share predicted states across trajectory candidates, thereby diluting the action-specific consequences that ought to guide selection. To bridge this gap, we propose DA-WAM, a framework that unifies predictive representation learning, action-conditioned future modeling, and trajectory scoring under a single decision-making objective. DA-WAM maintains predictive supervision throughout planner optimization via an online encoder and a stable momentum target, allowing future representations to co-evolve with the driving task. An action-conditioned predictor generates a distinct future latent state per trajectory candidate, which is then evaluated by a future-latent-conditioned factorized scorer. For the expert-matched trajectory, the predicted future latent is supervised by the observed future representation, while safety-critical hard negatives provide additional supervision near planning boundaries. Extensive experiments on NAVSIM-v1 and NAVSIM-v2 demonstrate state-of-the-art performance, while ablations and diagnostic analyses validate the key components.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑