MotionJEPA:通过在潜在空间中捕捉视觉变化来防止时间特征坍缩
MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space
浏览论文内容
中文总结 AI 辅助
MotionJEPA通过引入DISReg正则化器,在潜在空间中预测时间差分图像嵌入,平衡静态与动态特征学习,防止时间特征坍缩,从而提升下游规划成功率。
中文摘要 AI 辅助
联合嵌入预测架构(JEPAs)是一种无需视觉重建即可学习任务无关的潜在世界模型的有前景范式。然而,标准JEPA训练表现出对慢特征的强烈归纳偏置,导致特征抑制和潜在表征的坍缩。虽然逆动力学提供了时间上的抗坍缩能力,但它依赖于动作标签,并且对嵌入通用的、无标签的动力学几乎没有激励。我们引入了差分图像与单图像嵌入正则化(DISReg),这是一种新颖的正则化器,基于逆动力学风格的模块,无需任何像素重建损失即可预测时间差分图像嵌入,从而鼓励静态和动态特征的平衡学习。DISReg包含一个静态项,用于塑造图像嵌入的分布并鼓励慢特征,以及一个动态项,与对嵌入的直接正则化不同,它不对图像嵌入的形状或分布施加约束,而仅激励动态特征的存在。通过将此正则化器集成到标准JEPA中,我们建立了新架构MotionJEPA。潜在探针实验表明,MotionJEPA比其他方法产生更完整的表征,我们的轨迹分析显示,它保持了几何上简单且曲率较低的潜在嵌入。我们进一步表明,MotionJEPA在四个环境中,在静态背景干扰物下,提高了下游规划的成功率。
英文摘要
Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction. However, standard JEPA training exhibits a strong inductive bias towards slow features, causing feature suppression and the collapse of latent representation. While inverse dynamics provides temporal anti-collapse, it relies on action labels and offers little incentive to embed general, unlabeled dynamics. We introduce Difference Image and Single image embedding Regularization (DISReg), a novel regularizer that builds on an inverse-dynamics-style module that predicts temporal difference image embeddings without any pixel reconstruction loss, encouraging balanced static and dynamic feature learning. DISReg consists of a static term that shapes the distribution of the image embedding and encourages slow features, and a dynamic term, which, unlike direct regularization on the embedding, imposes no constraint on the image embedding's shape or distribution and instead only incentivizes that dynamic features be present. By integrating this regularizer into a standard JEPA, we establish our new architecture, MotionJEPA. Latent probing demonstrates that MotionJEPA produces more complete representations than other methods, and our trajectory analysis shows it maintains geometrically simple latent embeddings with low curvature. We further show that MotionJEPA improves downstream planning success under static-background distractors across four environments.
发表机构
- University of Oxford(牛津大学)
- vivo Tech Research GmbH(vivo科技研究有限公司)
- Bielefeld University(比勒费尔德大学)
- Slater Labs(Slater实验室)
- Cold Spring Harbor Laboratory(冷泉港实验室)
- Brown University(布朗大学)
- AMI Labs(AMI实验室)
- vivo BlueImage Lab, vivo Mobile Communication Co., Ltd., China(vivo蓝心影像实验室,维沃移动通信有限公司)
机构由 AI 辅助整理,请以论文原文为准。