发表机构
The Hong Kong University of Science and Technology (Guangzhou); COCO Matrix(香港科技大学(广州); COCO Matrix)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对JEPA世界模型前向预测未明确物理状态可识别性的问题,提出PSG-JEPA模型,通过训练阶段的两个grounding目标提升性能,在三个评估层面均优于现有基线。
AI 中文摘要
学习结构化且与控制相关的隐层表示仍是世界模型的关键挑战。近期基于JEPA的世界模型从观测序列中学习动作条件下的预测隐层动力学,但它们的前向预测目标未明确要求从单个隐层识别机器人中心物理状态、从隐层对识别状态变化的可靠性,这会限制下游规划和策略性能。我们提出PSG-JEPA,一种物理 grounding的JEPA世界模型,它在前向预测之外,通过两个互补的grounding目标塑造隐层空间:将单个隐层与机器人本体感知状态 grounding,将隐层对与多 horizon关节角变化 grounding。两个目标仅在训练阶段应用,推理架构和计算成本保持不变。为全面评估PSG-JEPA,我们在三个层面开展实验:(1) 通过探测评估隐层可识别性;(2) 基于冻结隐层的目标条件规划;(3) 在仿真和真实机器人上的策略学习。实验表明,PSG-JEPA在所有三个层面均持续优于最先进的隐层世界模型基线。
英文摘要
Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent dynamics from observation sequences. However, their forward-prediction objectives do not explicitly enforce reliable identifiability of robot-centric physical state from individual latents or state changes from latent pairs, which can limit downstream planning and policy performance. We propose PSG-JEPA, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes. Both objectives are applied only during training, leaving the inference architecture and computational cost unchanged. To comprehensively evaluate PSG-JEPA, we conduct experiments at three levels: (1) latent identifiability via probing, (2) goal-conditioned planning on frozen latents, and (3) policy learning in simulation and on a real robot. Experiments demonstrate that our PSG-JEPA consistently outperforms state-of-the-art latent world-model baselines at all three levels.