AI 中文总结
XP-JEPA通过将视觉与物理表征交叉预测,提升了潜在动力学的可预测性,在多任务套件上降低了展开漂移并提高了控制成功率,无需特权输入即可实现更优的基于展开的控制性能。
AI 中文摘要
潜在世界模型通过预测候选动作如何变换已学习的表征来进行规划。然而在自预测模型中,编码器与预测器是联合优化的,可能会共同适应那些易于预测但仅受场景物理演化弱约束的潜在转移。我们引入交叉预测JEPA(XP-JEPA),其将视觉潜在动力学建立在特权物理轨迹的基础上。XP-JEPA分别编码视觉观测与物理状态,通过共享的动作条件预测器推进两者,并将每个预测与两种模态的未来表征匹配。该目标鼓励两种模态间的统一潜在动力学,且基于底层物理转移。训练后物理分支被丢弃,部署时仅保留视觉模型。在涵盖六个评估子族的多任务套件上,XP-JEPA将新拟合预测器的展开漂移从0.361降至0.104,并将平均控制成功率从53.6%提升至78.2%。直接物理状态回归提高了位置可解码性,但使可预测性和控制性能接近仅视觉基线。因此,交叉预测物理基础可为基于展开的控制生成更具可预测性的潜在动力学,且测试时无需特权输入。
英文摘要
Latent world models plan by predicting how candidate actions advance learned latent dynamics. In self-predictive models, however, the encoder and predictor are optimized jointly and can co-adapt to latent transitions that are easy to predict but weakly constrained by the physical evolution of the scene. We introduce the cross-predictive JEPA (JEPA-x), which grounds latent dynamics in privileged physical trajectories. JEPA-x treats visual observations and physical states as corresponding views of the same action-conditioned trajectory, advances both through a shared predictor, and matches each prediction to the future representations of both modalities. This encourages the action-conditioned predictor to learn a common transition rule across the two views. Privileged physical state is used only during training, leaving a visual-only model at deployment. Empirical results show that JEPA-x reduces the rollout drift of a newly fitted predictor from $0.361$ to $0.104$ and increases mean control success from $53.6\%$ to $78.2\%$ on a multi-task suite spanning six evaluation subfamilies. We additionally show that direct physical-state regression improves decodability without improving forecastability or control, indicating that the benefit comes from shaping latent dynamics rather than merely encoding physical variables.
CommentsUnder review