Hamiltonian JEPA:具有继承控制状态的行动条件世界模型
Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
- The Blavatnik School of Computer Science(布拉瓦特尼克计算机科学学院)
- Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
H-JEPA通过分离感知与控制的联合嵌入预测架构,利用端口哈密顿动力学和端口逆一致性,在像素控制基准上超越基线,最大增益达12.6个百分点。
AI中文摘要:
从像素进行规划需要的不仅仅是一个稳定且可预测的潜在空间。规划器所评分的状态还必须按照行动如何移动系统来组织。联合嵌入预测架构(JEPAs)通过预测未来表示来避免像素重建,但现有的行动条件JEPAs要求一个嵌入同时服务于感知和控制。我们引入了H-JEPA,它将两者分离。一个宽感知码通过Bures-Wasserstein先验被正则化到良好缩放的各向同性几何,该码的一个固定正交归一化切片是控制状态,它继承了码的协方差而无需自身的目标。该状态在相位条件的耗散端口哈密顿动力学下演化,其输入端口具有正交归一化列。端口逆一致性(PIC)通过该端口的转置读回所执行的行动。我们证明这种读出恰好是投射到端口方向上的展开误差,因此PIC是预测误差的无参数重新加权,而非辅助行动解码器。将读出与端口分离会破坏这一恒等式并损失一半的增益。H-JEPA在四个基于像素的控制基准上,在至多10个训练周期后,匹配或超过了无重建基线(包括行动解码的Delta-JEPA),其最大增益出现在OGB-Cube上(91.9对79.3个百分点)。在PushT和OGB-Cube上的消融研究分离了结构化预测器、PIC、预测范围、状态秩和抗坍缩先验的贡献。
英文摘要:
Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing action-conditioned JEPAs ask one embedding to serve both perception and control. We introduce H-JEPA, which separates the two. A wide perceptual code is regularized toward a well-scaled isotropic geometry with a Bures-Wasserstein prior, and a fixed orthonormal slice of that code is the control state, which inherits the code's covariance without any objective of its own. The state evolves under phase-conditioned dissipative port-Hamiltonian dynamics whose input port has orthonormal columns. Port-inverse consistency (PIC) reads the executed action back through the transpose of that port. We show that this readout is exactly the rollout error projected onto the port directions, so PIC is a parameter-free reweighting of prediction error and not an auxiliary action decoder. Untying the readout from the port breaks this identity and loses half of the gain. H-JEPA matches or exceeds reconstruction-free baselines, including the action-decoding Delta-JEPA, on four pixel-based control benchmarks after at most $10$ training epochs, and its largest gain is on OGB-Cube ($91.9$ against $79.3$ percent). Ablations on PushT and OGB-Cube separate the contributions of the structured predictor, PIC, the prediction horizon, the state rank, and the anti-collapse prior.