发表机构
Central South University; Yanshan University(中南大学; 燕山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对JEPA世界模型丢弃动作因果动态信息的问题,提出动作锚定视觉不变潜在(AVL)方法,通过动作锚点和视觉不变路径防止信息坍缩,在四个机器人任务上显著提升扰动下的成功率。
AI 中文摘要
联合嵌入预测架构(JEPA)通过预测未来的潜在表示而不重建观测,使世界模型能够聚焦于高层语义动态。然而,JEPA可能在保留高维视觉信息的同时,丢弃关于动作物理后果的信息。我们将这种失败模式称为因果动态信息坍缩,并提出动作锚定的视觉不变潜在(AVL)方法来防止这种坍缩。我们首先将执行的动作作为辅助动态锚点,鼓励模型保留动态信息;然后使用视觉不变路径,在不丢弃动态信息的情况下对齐扰动和干净的潜在预测,迫使模型充分理解和利用因果动态信息。我们在四个机器人控制任务(TwoRoom、PushT、OGBench Cube和Reacher)上验证了AVL,结果表明在视觉扰动下它显著提高了成功率,同时保持了干净环境下的性能。我们进一步评估了物理后果对齐、干净-噪声动态一致性以及目标转移子空间擦除的因果效应。综合这些结果表明,在AVL下,与规划因果相关的动态信息得以保留而不发生坍缩。
英文摘要
Joint embedding predictive architectures (JEPAs) predict future latent representations without reconstructing observations, enabling world models to focus on high-level semantic dynamics. However, a JEPA can preserve high dimensional visual information while discarding information about the physical consequences of actions. We call this failure mode causal dynamics information collapse and propose action-grounded vision-invariance latent (AVL) to prevent this collapse. We first use the executed action as an auxiliary dynamics anchor that encourages the model to preserve dynamics information, and then use a vision-invariance pathway which aligns perturbed and clean latent predictions without discarding dynamics information, forcing the model to fully understand and utilize causal dynamics information. We validate AVL on four robotic control tasks (TwoRoom, PushT, OGBench Cube, and Reacher), showing that it substantially improves success rates under visual perturbations while preserving clean-environment performance. We further evaluate physical consequence alignment, clean-noisy dynamics consistency, and the causal effect of targeted transition subspace erasure. Collectively, these results indicate that dynamic information causally relevant to planning is preserved from collapse under AVL.
Comments17 pages, 10 figures, 12 tables