发表机构
Southeast University; NIO; Nanyang Technological University; Tongji University(东南大学; 蔚来; 南洋理工大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PlanWAM通过规划任务塑造未来表征,结合时间寄存器金字塔、特权未来后验分支与知识蒸馏,在NAVSIM、HUGSIM数据集上实现领先的自动驾驶规划性能。
AI 中文摘要
端到端自动驾驶中的世界模型会预测未来场景演化,为轨迹规划提供前瞻性。现有方法主要研究如何预测未来以及如何利用未来,但较少关注哪种未来表征对规划最有用。为此,我们提出PlanWAM,即规划塑造型世界动作模型,核心思路是让规划任务塑造未来状态表征,使其保留对规划最有用的信息。潜在世界模型随后从历史观测中预测该规划塑造型未来潜在表征,并将其用于规划,实现前瞻性规划。具体而言,我们首先使用时间寄存器金字塔以感知近期的方式压缩多帧历史信息,学习面向未来推理和规划的紧凑历史表征;接着引入特权未来后验分支,观测真实未来帧,并用轨迹规划目标塑造其未来潜在表征,得到规划塑造型未来潜在表征;通过事后到前瞻的知识蒸馏训练仅依赖历史的先验分支来预测该未来潜在表征,预测的未来潜在表征作为规划上下文,指导轨迹生成与选择。PlanWAM在NAVSIM-v1/v2导航测试中达到93.8 PDMS/90.9 EPDMS,在零样本设置下的闭环HUGSIM中达到38.7 HD-Score,在开环和闭环评估中均展现出领先的规划性能,大量实验进一步证明,规划塑造型未来表征为世界动作模型提供了有效且可部署的前瞻性形式。
英文摘要
World models in end-to-end autonomous driving predict future scene evolution to provide foresight for trajectory planning. Existing methods mainly study how to predict the future and how to use it, but less often ask which future representation is actually most useful for planning. To this end, we propose PlanWAM, a Planning-Shaped World Action Model. The key idea is to let the planning task shape the future-state representation, so that it retains the information most useful for planning. A latent world model then predicts this planning-shaped future latent representation from historical observations and uses it for planning, enabling foresighted planning. Specifically, we first use a Temporal Register Pyramid to compress multi-frame historical information in a recency-aware manner, learning a compact history representation oriented toward future reasoning and planning. We then introduce a privileged future posterior branch that observes ground-truth future frames, and shape its future latent representation with trajectory-planning objectives to obtain a planning-shaped future latent representation. Hindsight-to-Foresight Distillation trains a prior branch that depends only on history to predict this future latent representation. The predicted future latent representation serves as planning context and guides trajectory generation and selection. PlanWAM achieves 93.8 PDMS / 90.9 EPDMS on NAVSIM-v1/v2 navtest and reaches 38.7 HD-Score on closed-loop HUGSIM in a zero-shot setting, demonstrating leading planning performance across both open-loop and closed-loop evaluations. Extensive experiments further demonstrate that planning-shaped future representations provide an effective and deployable form of foresight for world-action models.