发表机构
Huazhong University of Science and Technology; Shanghai Zaofu Intelligent Technology Co., Ltd.(华中科技大学; 上海造孚智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ForeDrive提出一种规划相关的潜在世界模型,通过非对称耦合扩散Transformer规划器,利用多时间范围未来潜在表示作为引导,实现纯模仿学习下的高性能端到端自动驾驶。
AI 中文摘要
现有的潜在世界模型通常针对未来可预测性进行优化,但由此产生的表示不一定对自动驾驶中的规划有用。预测通常用于预训练或辅助监督,而非作为轨迹生成的直接条件信号。我们提出ForeDrive,它学习一种与规划相关的潜在表示,并将其与非对称耦合的扩散Transformer(DiT)规划器相结合。规划器消费由JEPA风格世界模型学习到的多时间范围潜在未来表示;规划梯度更新共享的在线编码器,而停止梯度路由仅使用预测损失训练潜在预测器。由于预测的未来在不同时间范围上的可靠性不同,且BEV轨迹与图像标记不对齐,我们使用门控视觉融合、未来状态注入和轨迹自适应偏置(TAB)将未来潜在信息作为引导注入,而不覆盖当前观测。仅使用纯模仿学习训练,推理时仅使用当前前视图像作为视觉输入,ForeDrive在NAVSIM v1上达到89.9 PDMS,在NAVSIM v2上达到90.0单阶段EPDMS,无需强化学习或外部轨迹评分器。
英文摘要
Existing latent world models are typically optimized for future predictability, yet the resulting representations are not necessarily useful for planning in autonomous driving. Predictions are commonly used for pretraining or auxiliary supervision rather than as direct conditioning signals for trajectory generation. We propose ForeDrive, which learns a planning-relevant latent representation and couples it asymmetrically to a Diffusion Transformer (DiT) planner. The planner consumes multi-horizon latent future representations learned with a JEPA-style world model; planning gradients update the shared online encoder, while stop-gradient routing trains the latent predictor with forecasting losses only. Because predicted futures have varying reliability across horizons and BEV trajectories are misaligned with image tokens, we use gated visual fusion, future-status injection, and Trajectory-Adaptive Bias (TAB) to inject future latents as guidance without overriding the current observation. Trained with pure imitation learning and using only the current front-view image as visual input at inference, ForeDrive attains 89.9 PDMS on NAVSIM v1 and 90.0 one-stage EPDMS on NAVSIM v2, without reinforcement learning or an external trajectory scorer.
Comments9 pages, 4 figures; 8 pages supplementary with 4 figures