发表机构
Yale University; Jump Trading; Brown University(耶鲁大学; 跃动交易公司; 布朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Semigroup-JEPA,通过动作条件注入物理参数并联合训练编码器与预测器,实现零样本物理泛化,显著降低预测误差并提升控制成功率。
AI 中文摘要
联合嵌入预测架构(JEPA)世界模型学习世界的紧凑潜在表示,以支持预测和规划,但其学习物理并生成物理上真实动力学的能力迄今尚未得到测试。在这项工作中,我们引入了Semigroup-JEPA(SG-JEPA),它扩展了LeWorldModel框架,通过动作条件将控制物理的参数提供给时间模型,并通过自回归潜在展开联合训练编码器和预测器。为了评估模型在分布外泛化的能力,我们设计了不同引力场下的动力学任务,这些任务尽管遵循相同的物理定律,却表现出定性不同的动力学,从弱引力场中的漂浮运动到强引力场中的快速弹跳。与DINO-WM相比,SG-JEPA在二维数据集上将开环预测误差降低了最多2倍,并在三维机器人数据集上将控制成功率提高了最多2.5倍,为此我们训练了独立的扩散策略。为了解释这一优势,我们开发了一个线性特征模型,将局部定律条件误差与其在展开下的递归放大分离开来。在该模型的指导下,我们发现将多步展开损失反向传播到表示中,训练编码器保留预测器能够前向传播的特征,而这些特征正是动力学所依赖的,因此大部分增益来自编码器学习更好的特征,而非预测器学习更好的动力学。参见项目页面:此https URL。
英文摘要
Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning and jointly training an encoder and predictor through an autoregressive latent rollout. To evaluate the model's ability to generalize out of distribution, we design dynamical tasks under different gravitational fields that, despite obeying the same physical law, exhibit qualitatively different dynamics, ranging from floating motion in weak gravitational fields to rapid bouncing in strong ones. In contrast to DINO-WM, SG-JEPA reduces open-loop prediction error by up to 2 times on two-dimensional datasets, and increases control success rate up to 2.5 times for three-dimensional robotic datasets, for which we train independent diffusion policies. To explain this advantage, we develop a linear feature model that separates local law-conditioned error from its recursive amplification under rollout. Guided by this model, we find that back-propagating the multi-step rollout loss into the representation trains the encoder to keep the features that the predictor can carry forward, and that those are the features the dynamics depend on, so most of the gain comes from the encoder learning better features rather than from the predictor learning better dynamics. See project page at https://sg-jepa.github.io.