发表机构
Honda Research Institute USA; Johns Hopkins University(本田美国研究院; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对测试时动态变化导致世界模型预测不可靠的问题,提出JEPA-TTT方法,通过持续自监督测试时训练调整潜在动态预测器,在多种动态变化下显著降低预测误差并提升规划性能。
AI 中文摘要
世界模型使智能体能够通过预测环境的未来状态来进行规划,但当测试时的动态与训练时所见不同时,其预测可能变得不可靠。我们提出了JEPA-TTT,它在整个测试时间内调整预训练的、基于动作的条件联合嵌入预测架构(Joint-Embedding Predictive Architecture)世界模型的潜在动态预测器。自监督更新跨回合累积,而视觉编码器和奖励头保持固定,从而保留预训练的表示和任务目标。规划既不需要目标图像,也不需要在线环境奖励。JEPA-TTT使用密集回放,在每个时间偏移处形成预测窗口,将其保留在不断增长的缓冲区中,并从该缓冲区采样小批量用于预测器更新。在四个连续控制环境中的八种动态变化下,JEPA-TTT在每种变化上都改善了规划性能。经过500个测试时回合后,它将自回归潜在预测误差平均降低了83%,并将规划性能相对于冻结的JEPA世界模型提高了153%。这些结果表明,持续的自我监督测试时训练可以在动态变化下调整预训练的潜在世界模型。
英文摘要
World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time. Self-supervised updates accumulate across episodes, while the visual encoder and reward head remain fixed, preserving the pretrained representation and task objective. Planning requires neither a goal image nor online environment reward. JEPA-TTT uses dense replay, which forms prediction windows at every temporal offset, retains them in a growing buffer, and samples minibatches from that buffer for predictor updates. Across eight dynamics shifts in four continuous-control environments, JEPA-TTT improves planning on every shift. After 500 test-time episodes, it reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model. These results show that persistent self-supervised test-time training can adapt a pretrained latent world model under changed dynamics.
CommentsAccepted to World Models in Physical AI Workshop @ NeurIPS 2026 | Project page: https://jepa-ttt.github.io/