AI 中文总结
Auto-JEPA是面向端到端自动驾驶的连续意图隐式世界模型,通过联合嵌入预测学习未来驾驶意图,无需密集未来世界建模,在NAVSIM数据集上取得优异规划性能,可聚焦规划相关视觉特征。
AI 中文摘要
现有自动驾驶世界模型通常对未来视频、占据状态、鸟瞰图(BEV)表示或智能体运动进行密集预测。本文认为规划无需重构完整未来世界,仅需关注影响未来自车动作的场景特征。基于此,我们提出Auto-JEPA,这是一种面向动作的隐式世界模型,通过联合嵌入预测学习连续未来驾驶意图。给定视觉观测、自运动历史和导航指令,Auto-JEPA预测与未来自车轨迹隐式表示对齐的意图嵌入;该预测意图从固定轨迹记忆中检索可执行轨迹,再由场景条件候选选择模块排序。Auto-JEPA保持视觉编码器冻结,无需显式感知标注,也不使用学习到的轨迹生成器;仅优化轨迹表示、意图预测和候选选择等任务特定模块,在NAVSIM v1上达到91.3的PDMS,在NAVSIM v2上达到89.1的EPDMS。语义遮挡实验显示,遮蔽动态智能体区域引发的平均意图变化是等面积随机遮蔽的2.97倍;遮蔽对未来驾驶有实质影响的车辆会显著改变预测意图和选定轨迹,而遮蔽无影响车辆时两者基本不变。这些结果表明,未来意图预测促使模型聚焦规划相关视觉特征,无需密集未来世界建模即可支持高质量规划。
英文摘要
Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion. We argue that planning need not reconstruct the complete future world, but only focus on scene features that affect future ego action. Based on this perspective, we propose Auto-JEPA, an action-oriented latent world model that learns continuous future driving intent through joint-embedding prediction. Given visual observations, egomotion history, and navigation commands, Auto-JEPA predicts an intent embedding aligned with the latent representation of the future ego trajectory. The predicted intent retrieves executable trajectories from a fixed trajectory memory, which are then ranked by a scene-conditioned candidate selection module. Auto-JEPA keeps the visual encoder frozen, requires no explicit perception annotations, and uses no learned trajectory generator. By optimizing only task-specific modules for trajectory representation, intent prediction, and candidate selection, Auto-JEPA achieves 91.3 PDMS on NAVSIM v1 and 89.1 EPDMS on NAVSIM v2. Semantic occlusion experiments show that masking dynamic-agent regions induces an average intent change 2.97x that of equal-area random masking. Moreover, occluding vehicles that affect future driving substantially changes the predicted intent and selected trajectory, whereas both remain essentially unchanged when non-influential vehicles are occluded. These results show that future-intent prediction encourages the model to focus on planning-relevant visual features and supports high-quality planning without dense future-world modeling.