arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向物理接地的JEPA世界模型,用于目标条件机器人规划

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

Muyuan Liu, Yue Huang, Zheng Liang, Xiang Gao

arXiv 2609.03565首次发表:更新:

发表机构

GENISOM AI(GENISOM AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出结合逆动力学模型与状态对齐的端到端JEPA世界模型,在四个机器人规划基准任务中多数达到最优成功率,验证了状态对齐对提升规划效果的作用。

AI 中文摘要

基于动作条件的JEPA世界模型可在不重构未来像素的情况下,朝着视觉指定的目标进行规划,但仅潜在预测无法明确鼓励学习到的表征保留与机器人控制相关的信息。我们提出一种端到端的JEPA世界模型,该模型通过逆动力学模型(IDM)和状态对齐(SA)增强潜在预测。逆动力学模型可防止潜在崩溃,使潜在转换能反映产生该转换的动作;状态对齐则将连续的表征与其关联的物理配置和运动绑定。在四个基准任务中,我们的模型在TwoRoom任务上达到100%的成功率,在PushT任务上达到98%的成功率,在OGBench-Cube任务上达到87%的成功率,在Reacher任务上的表现与LeWorldModel相当。我们的 ablation 实验进一步表明,在所有四个任务中,添加状态对齐均比单独使用IDM持续提升规划成功率。尽管我们的主要基线LeWorldModel在OGBench-Cube任务上的平均拉直程度更高,但转换子空间分析显示,其转换能量集中在维度显著更低的子空间中。我们的状态对齐模型的有效转换维度高于LeWorldModel,且比单独使用IDM更能提升规划效果,这表明状态对齐是逆动力学模型用于机器人规划的有效补充。

英文摘要

Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and makes latent transitions informative of the actions that produced them, state alignment grounds consecutive representations in their associated physical configuration and motion. Across four benchmark tasks, our model attains the highest success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to LeWorldModel on Reacher. Our ablation further shows that adding state alignment consistently improves planning success over IDM alone across all four tasks. Although LeWorldModel, our primary baseline, attains higher average straightening on OGBench-Cube, transition-subspace analysis shows that its transition energy is concentrated in a substantially lower-dimensional subspace. Our state-aligned model exhibits a higher effective transition dimension than LeWorldModel and improves planning over IDM alone, supporting state alignment as an effective complement to inverse dynamics for robotic planning.

Comments5 pages, 4 figures, 2 tables. Accepted to the IROS 2026 Workshop on Physical World Models for Scaling Embodied AI (PWMS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑