发表机构
School of Artificial Intelligence, Beihang University; Shanghai Jiao Tong University; Dim12 AI; GigaAI; National University of Singapore; TeleAI(北京航空航天大学人工智能学院; 上海交通大学; Dim12 AI; 极佳科技; 新加坡国立大学; TeleAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Drive-HWM是一种分层快慢世界建模框架,通过动态感知隐变量引导,在NAVSIM数据集上实现了出色的自动驾驶性能,验证了其设计的有效性。
AI 中文摘要
世界模型为自动驾驶提供了一种有前景的范式,通过预测交通场景的演化并利用此类预测支持动作生成。然而,现有方法要么将未来预测与动作生成分开,要么在相同时间尺度上联合预测,难以同时实现长时程预测和基于观测的响应式决策。我们提出Drive-HWM,一种分层快慢世界建模框架,在互补时间尺度上组织未来表示预测和动作生成。慢世界模型预测多步未来表示以捕捉扩展场景演化。为显式建模驾驶环境中丰富的运动动态,我们引入通过光流预测学习的动态感知隐变量。在这些未来表示的引导下,快模型使用轻量多模态骨干网络和自回归专家,从最新观测中联合预测下一帧和即时动作。下一帧预测促使快模型捕捉临近场景演化,而单步动作生成允许决策随新观测的到达持续更新。在NAVSIM v1和v2上的大量实验证明了Drive-HWM的强大驾驶性能。全面的 ablation 研究进一步验证了分层快慢设计、动态感知未来表示以及联合下一帧与动作预测的有效性。
英文摘要
World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and using such predictions to support action generation. However, existing approaches either separate future prediction from action generation or jointly predict them at the same temporal scale, making it difficult to simultaneously achieve long-horizon anticipation and responsive, observation-grounded decision making. We present Drive-HWM, a hierarchical slow--fast world modeling framework that organizes future representation prediction and action generation at complementary temporal scales. The slow world model predicts multi-step future representations to capture extended scene evolution. To explicitly model the abundant motion dynamics in driving environments, we introduce Dynamic-Aware Latents learned through optical-flow prediction. Guided by these future representations, the fast model uses a lightweight multimodal backbone and an autoregressive expert to jointly predict the next frame and the immediate action from the latest observation. Next-frame prediction encourages the fast model to capture imminent scene evolution, while one-step action generation allows decisions to be continuously updated as new observations arrive. Extensive experiments on NAVSIM v1 and v2 demonstrate the strong driving performance of Drive-HWM. Comprehensive ablation studies further validate the effectiveness of the hierarchical slow--fast design, dynamics-aware future representations, and joint next-frame and action prediction.
Comments14 pages