arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JEPA-WAM:基于联合嵌入世界建模的视觉-语言-动作策略学习

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Yihan Lin, Jiawei He, Shifeng Bao, Chen Zhao, Yang Li, Xiaobo Wang, Yan Wang, Cheng Chi, Jing Zhang

arXiv 2608.09381首次发表:更新:

发表机构

XYZ Embodied AI(XYZ具身智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出JEPA-WAM,一种构建于预训练V-JEPA空间的隐式WAM,通过共享预测器耦合隐式转移预测与动作生成,在LIBERO-Plus等数据集上取得优异的机器人操作策略性能,且泛化能力强。

AI 中文摘要

鲁棒的机器人控制得益于对状态转移的显式建模,但视频生成式世界动作模型(WAMs)会带来可观的部署成本。现有的隐式WAMs避免了显式的未来生成,却常常压缩预测表征,或将预测建模与用于动作生成的表征相分离。我们提出JEPA-WAM,一种构建于预训练V-JEPA空间中的隐式WAM,它通过共享预测器将隐式转移预测与连续动作生成耦合起来。JEPA-WAM预测空间结构化的联合当前-未来目标,该目标捕获当前与未来观测间任务共享的视觉时间结构,同时保留密集的补丁级对应关系。通过共享预测器,转移监督直接塑造主干网络,从中提取专用表征用于动作预测。该设计也可在预训练的VLA策略中实例化,同时保留其原有的感知与动作通路。在LIBERO-Plus上,JEPA-WAM取得79.2%的成绩,是未进行大规模机器人策略预训练的最佳结果;其预训练π₀.5实例达到86.3%,实现了最佳整体性能。在RoboTwin 2.0及真实世界双手机械臂操作上的实验进一步证明,其在视觉与空间偏移下具备强泛化能力。

英文摘要

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained $π_{0.5}$ instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.

Comments22 pages, 7 figures. Project page: https://spritewithoutice.github.io/JEPA_WAM/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑