arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AD-E2E-JEPA:一种用于端到端自动驾驶的联合嵌入预测架构

AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving

Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska

arXiv 2609.34085首次发表:更新:

发表机构

New York University; AMI Labs(纽约大学; AMI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AD-E2E-JEPA,通过SIGReg正则化可学习投影器大幅压缩规划补丁和嵌入维度,实现100倍推理加速并保持规划性能,在零样本规划中有效,且预训练投影器提升下游模仿学习性能。

AI 中文摘要

自动驾驶需要能够理解物理世界、进行推理和规划并安全运行的世界模型。在本文中,我们首先系统地评估了现有的用于端到端自动驾驶(E2EAD)的动作条件联合嵌入预测架构(JEPA)世界模型,包括LeWM、DINO-WM和JEPA-WM。为了将世界模型的质量与策略学习分离,我们采用了一种目标条件零样本规划设置,该设置使用真实未来观测作为目标来评估这些模型,而无需训练任何驾驶策略。我们发现,现有的基于JEPA的世界模型要么对驾驶准确但计算成本高,要么计算高效但不足以进行规划。为了解决这一权衡,我们提出了AD-E2E-JEPA,它引入了一个经过SIGReg正则化的可学习投影器,应用于投影后的补丁嵌入。该投影器将规划补丁数量减少了16倍,嵌入维度减少了4倍,实现了100倍的推理加速,同时保持了规划性能,在256个候选轨迹上执行8帧推演仅需0.8秒。在不训练任何驾驶策略的情况下,世界模型本身平均能达到位于20米外的目标,在分别包含256和8192个候选轨迹的轨迹词汇表上,推演位移分别为4.0米和2.8米。在NAVSIMv2基准上,在目标条件零样本规划中,使用乘法安全指标时EPDMS为67.3/72.9,不使用安全指标时EPDMS†为84.1/86.5。实验进一步表明,自监督预训练的投影器将下游模仿学习性能从80.2提高到85.4 EPDMS。源代码可在以下网址获取:https URL

英文摘要

Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose \textbf{AD-E2E-JEPA}, which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by $16\times$ and the embedding dimension by $4\times$, achieving a $100\times$ inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. \textit{Without} training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/2.8 meters, using world-model rollouts over trajectory vocabularies of respectively 256/8,192 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/72.9 EPDMS with multiplicative safety metrics and 84.1/86.5 EPDMS$^{\dagger}$ without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS. The source code is available at https://github.com/HaoranZhuExplorer/AD-E2E-JEPA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑