发表机构
Tsinghua University; Institute of Automation, Chinese Academy of Sciences; Northeastern University; Harvard University; University of Oxford(清华大学; 中国科学院自动化研究所; 东北大学; 哈佛大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文揭示JEPA世界模型因单帧渲染而缺乏速度信息,提出TI-JEPA通过分离姿态码和运动码修复,在多个基准上显著提升规划性能。
AI 中文摘要
用于世界建模的联合嵌入预测架构(JEPAs)训练一个编码器,使得预测器能够从当前嵌入和动作映射到下一帧的嵌入,且始终基于单个渲染帧。这存在一个结构性的盲点:没有运动模糊的渲染器仅根据配置绘制场景,因此单帧嵌入不携带任何速度信息,对于任何编码器(包括官方发布的LeWM权重)都是如此。我们在四个真实基准(PushT、Reacher、Cube、TwoRoom)的官方检查点上证实了这一点:每个线性速度探针的准确率都处于或低于随机水平,而位置探针的R²达到约0.95。我们引入了RateIdent,一个三阶段的诊断协议,以及TI-JEPA,一个轻量级的修复方案,将潜在变量分为姿态码和显式的有限差分运动码,并联合预测。在三个物理基础环境中,TI-JEPA在停止目标规划任务上相比匹配内存的基线获得了显著的、对种子稳健的提升,例如在Pendulum上最终距离降低55%(p=3.2x10^-10),在CartPole上降低64%(p=5.1x10^-15)。我们在官方ViT-Tiny加AdaLN-transformer规模上复现了这一结果,然后将相同的方案应用于从零开始训练的真实dm_control Reacher照片,其中TI-JEPA的分支分离度比具有内存的基线高出约38倍,这是论文中最大的差距。与同规模循环RSSM风格预测器相比,TI-JEPA在三个环境中的两个上匹配或超越了其 rollout 准确率,保持姿态和运动的可分别探测性,并在耦合最强的环境中完全胜出。一个可验证的形式论证和六个评估环境表明,当速度重要时,单帧目标不是正确的预测对象,而一个小的、可解释的结构性改变可以在没有特权监督的情况下修复这一问题。代码、检查点和项目页面链接在标题下方。
英文摘要
Joint-embedding predictive architectures (JEPAs) for world modeling train an encoder so a predictor maps a current embedding and action to the next frame's embedding, always from a single rendered frame. This has a structural blind spot: a renderer without motion blur draws a scene from configuration alone, so a single-frame embedding carries no velocity information, for any encoder, including the official released LeWM weights. We confirm this on official checkpoints across four real benchmarks (PushT, Reacher, Cube, TwoRoom): every linear velocity probe sits at or below chance while position probes reach R^2 about 0.95. We introduce RateIdent, a three-stage diagnostic protocol, and TI-JEPA, a lightweight fix splitting the latent into a pose code and an explicit finite-difference motion code, predicted jointly. Across three physically grounded environments, TI-JEPA gives a significant, seed-robust gain on a stop-at-goal planning task over a matched-memory baseline, e.g. 55% lower final distance on Pendulum (p=3.2x10^-10) and 64% on CartPole (p=5.1x10^-15). We reproduce this at official ViT-Tiny plus AdaLN-transformer scale, then push the same recipe onto real dm_control Reacher photographs trained from scratch, where TI-JEPA's branch separation exceeds the memory-having baseline's by roughly 38x, the paper's largest margin. Against a same-footprint recurrent RSSM-style predictor, TI-JEPA matches or beats its rollout accuracy on two of three environments, stays separately probeable for pose and motion, and wins outright on the most coupled one. A checkable formal argument and six evaluated environments show single-frame targets are the wrong object to predict when velocity matters, and a small, interpretable structural change fixes it with no privileged supervision. Code, checkpoints, and the project page are linked below the title.
Comments32 pages, 17 figures, 12 tables. Code, model checkpoints, and project page are available via links in the paper