arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需高斯:用于JEPA世界模型的对比逆动力学

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

Jack Boylan, Chris Hokamp

arXiv 2608.17542首次发表:更新:

发表机构

Quantexa(昆泰克萨公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出AC-MTM,用对比逆动力学替代JEPA世界模型的高斯型抗坍塌机制,在多目标视觉任务上性能优于SIGReg,且训练稳定无额外复杂组件

AI 中文摘要

联合嵌入预测架构(JEPA)通过预测未来嵌入来学习世界模型,但该目标存在恒定编码器的平凡解,因此每个实际系统都会添加抗坍塌机制(LeCun, 2022; Assran等人, 2023; Bardes等人, 2022; 2024)。LeWorldModel(LeWM)通过SIGReg防止坍塌,SIGReg是一种正则化器,强制潜在分布匹配各向同性高斯:该表示通过规定其必须呈现的样子来稳定,与其建模的环境无关。我们认为抗坍塌压力可来自转移数据本身。动作对比掩码转移建模(AC-MTM)保留LeWM的前向潜在预测目标,并添加仅训练用的逆动力学头,该头通过Action-NCE训练:每个潜在转移必须在批次中的其他动作中识别产生它的动作,坍塌编码器可证无法完成该判别任务。逆分支在训练后丢弃,使测试时的编码、前向预测、规划和计算与LeWM完全相同。在匹配规划协议下的四个标准像素控制任务上,AC-MTM从头开始训练稳定,平均性能与SIGReg相当。在更难的多目标OGBench视觉场景任务上,结果与规定几何成为瓶颈一致:AC-MTM达到80.0±2.0%的成功率,而SIGReg为58.0±2.0%,每个训练种子提升20-24个百分点。单次50回合随机策略运行给出52%的基线估计。因此,对比逆动力学提供了无分布的抗坍塌信号,无需目标网络、停止梯度、预训练编码器或重建目标,且我们刻画了其成立的动作空间和可观测性假设。我们的代码可在该httpsURL获取

英文摘要

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mechanism (LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024). LeWorldModel (LeWM) prevents collapse with SIGReg, a regularizer that forces the latent distribution to match an isotropic Gaussian: the representation is stabilized by prescribing what it must look like, independently of the environment it models. We argue that the anti-collapse pressure can instead come from the transition data itself. Action-Contrastive Masked Transition Modeling (AC-MTM) keeps LeWM's forward latent-prediction objective and adds a training-only inverse-dynamics head trained with Action-NCE: each latent transition must identify the action that produced it among the other actions in the batch, a discrimination task that a collapsed encoder provably fails. The inverse branch is discarded after training, leaving test-time encoding, forward prediction, planning, and compute identical to LeWM. On four standard pixel-control tasks under a matched planning protocol, AC-MTM trains stably from scratch and matches SIGReg on average. On the harder multi-object OGBench Visual Scene task, results are consistent with the prescribed geometry becoming a bottleneck: AC-MTM reaches 80.0$\pm$2.0% success versus 58.0$\pm$2.0% for SIGReg, improving by 20-24 points in each training seed. A single 50-episode random-policy run gives a 52% baseline estimate. Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and we characterize the action-space and observability assumptions under which it holds. We make our code available at https://github.com/jackboyla/action-contrastive-jepa

Comments17 pages, 5 figures. Code: https://github.com/jackboyla/action-contrastive-jepa

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑