arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhyLatent:为JEPA世界模型学习与动力学相关的表征

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

Xi Zeng, Haojie Ren, Ziying Song, Yuanbo Nie, Ross Drummond

arXiv 2608.05720首次发表:更新:

发表机构

School of Mechanical and Aerospace Engineering, Nanyang Technological University; School of Artificial Intelligence (School of Software), Yanshan University; The University of Sheffield(南洋理工大学机械与航天工程学院; 燕山大学人工智能学院(软件学院); 谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhyLatent为JEPA世界模型设计训练目标,解决其三类失效模式,在OGBench-Cube等数据集上显著降低失效率、提升MPC成功率,证明全局非坍塌不足以学习可靠的JEPA状态空间。

AI 中文摘要

我们提出PhyLatent,这是一种针对联合嵌入预测架构(JointEmbedding Predictive Architecture,JEPA)世界模型的、与动力学相关的训练目标。我们的核心发现是,防止全局潜在空间坍塌并不能确保表征保留物理状态与动作后果。我们识别出JEPA世界模型的三种失效模式:物理不变性坍塌、物理可识别性坍塌和反事实动力学坍塌。PhyLatent通过三条训练路径解决这些问题,分别是物理不变性、物理可识别性和反事实动力学,这些路径通过物理状态 grounding、未来表征对齐、静态视觉不变性、反事实分支分离和潜在去噪实现。在OGBench-Cube上,PhyLatent将三种失效率从15.60%、6.71%和8.41%分别降至7.53%、0.95%和4.62%,并将模型预测控制(Model Predictive Control,MPC)的成功率从70.0%提升至78.1%。在相同架构和规划器下,它还将TwoRooms上的成功率从81.0%提升至98.0%,并在Reacher和PushT上保持竞争力。这些结果表明,仅全局非坍塌不足以学习可靠的JEPA世界模型状态空间。

英文摘要

We propose PhyLatent, a dynamics-relevant training objective for Joint-Embedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not necessarily ensure that the learned representation preserves physically meaningful state and action relationships. We identify three failure modes: sensitivity to appearance changes that leave the physical state unchanged, insufficient separation of distinct physical states, and insufficient separation of different action-conditioned futures. We refer to these as Physical Invariance Collapse, Physical Distinguishability Collapse, and Counterfactual Dynamics Collapse, respectively. PhyLatent targets these failures through three coordinated training pathways, implemented with static visual invariance, physical state grounding, future representation alignment, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three collapse rates by 43.9%, 27.9%, and 47.1%, respectively, while improving model predictive control (MPC) success by 12.0 percentage points (17.2% relative). Across four visual-control tasks, average planning success increases by 6.62 percentage points (8.3% relative). These results show that global non-collapse alone is insufficient for learning a reliable JEPA world-model state space, and that explicitly preserving dynamics-relevant structure can improve closed-loop planning.

Comments26 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑