arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20974cs.CVcs.AI

WA-JEPA:重新思考面向自动驾驶中世界-动作建模的视频JEPA范式

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对V-JEPA不适用于自动驾驶规划的问题,提出WA-JEPA模型,采用混合未来掩码预训练与条件流匹配,在NAVSIM和HUGSIM基准上取得优于基线的规划性能。

中文摘要 AI 辅助

视频联合嵌入预测架构(V-JEPA)通过自监督隐特征预测从视频中学习强大的时空表示。然而,V-JEPA基于随机掩码补全和确定性回归构建,根本不适合自动驾驶规划,自动驾驶规划需要与动作紧密耦合的未来导向预测。为解决该问题,我们重新思考V-JEPA范式,提出WA-JEPA,一种专为自动驾驶规划设计的原生V-JEPA世界-动作模型。WA-JEPA不采用随机时空掩码,而是采用混合未来掩码预训练,模型从观测上下文推断未来隐变量。不同于确定性回归,我们将未来预测重构为对未来隐变量的条件流匹配,大幅提升模型为下游规划生成合理未来隐变量的能力。最后,提出联合未来-动作预测器,在统一时空隐空间中共同去噪未来场景令牌和自车轨迹,使动作监督能直接塑造与规划相关的世界表示。WA-JEPA在nuPlan视频上预训练并在NAVSIM上微调,在NAVSIM-v2上达到91.7 EPDMS,超过最强的端到端和世界-动作基线1.6和1.3 EPDMS;且未进行HUGSIM特定微调时,在相同评估协议下的闭环HUGSIM基准上取得0.4462的最佳HD-Score。这些结果验证了原生V-JEPA世界-动作建模是一种强大且可扩展的自动驾驶规划范式。代码可在this https URL获取。

英文摘要

Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion and deterministic regression, making it fundamentally ill-suited for autonomous driving planning that demands future-directed prediction tightly coupled with action. To address this, we rethink the V-JEPA paradigm and present WA-JEPA, a V-JEPA-native world-action model designed for autonomous driving planning. Instead of random spatiotemporal masking, WA-JEPA employs hybrid future-masked pre-training, where the model infers future latents from observed context. Departing from deterministic regression, we recast future prediction as conditional flow matching over latent futures, which substantially improves the model's ability to generate plausible future latents for downstream planning. Finally, a joint future-action predictor is proposed to denoise future scene tokens and ego trajectories together in a unified spatiotemporal latent space, allowing action supervision to directly shape planning-relevant world representations. Pre-trained on nuPlan videos and fine-tuned on NAVSIM, WA-JEPA reaches 91.7 EPDMS on NAVSIM-v2, surpassing the strongest end-to-end and world-action baselines by 1.6 and 1.3 EPDMS, and, without HUGSIM-specific fine-tuning, attains the best HD-Score of 0.4462 on the closed-loop HUGSIM benchmark under the same evaluation protocol. These results validate V-JEPA-native world-action modeling as a powerful and scalable paradigm for autonomous driving planning. Code is available at https://github.com/AFARI-Research/WA-JEPA.

↑