arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31161cs.LG

我行动故我在:JEPA的动作条件化何时足以学习因果机制?

I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

发表机构澳大利亚负责任人工智能研究中心 · 阿德莱德大学澳大利亚机器学习研究所
查看机构详情
  • Responsible AI Research Centre, Australia(澳大利亚负责任人工智能研究中心)
  • Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究JEPA架构在动作条件下恢复潜在因果状态的条件,提出信息论目标与A-JEPA模型,并验证其可辨识性与泛化能力。

中文摘要 AI 辅助

近期的实证和理论进展表明,联合嵌入预测架构(JEPAs)可能为动作条件下的未来结果预测学习有意义的表征,从而成为世界模型的基础结构之一。然而,一般而言,准确的预测并不必然意味着恢复产生观测动态的潜在因果状态。本研究探讨了JEPAs在何种情况下以及如何从观测中恢复潜在因果状态。我们首先引入一个潜在变量模型,其中高维观测由潜在因果状态生成,其动态由动作条件化的转移机制控制。基于这一表述,我们开发了一个通用的信息论目标,该目标结合了用于学习转移动态的条件似然最大化和用于保留潜在状态信息的熵最大化。随后,我们建立了可辨识性条件,在这些条件下,由该通用目标学习的表征能够恢复潜在因果状态,直至分量可逆变换和置换。实现这种可辨识性的一个关键条件是转移机制中足够的动作诱导变化。在此发现的指导下,我们用动作调制的加性高斯噪声模型实例化该通用目标,产生动作调制JEPA(A-JEPA)。在合成环境上的实验验证了在可辨识性条件下的理论发现以及对中等程度违反的鲁棒性,而视觉基准则展示了改进的状态恢复和对未见转移机制的泛化能力。

英文摘要

Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models. However, accurate prediction does not, in general, necessarily imply recovery of underlying causal states that give rise to the observed dynamics. This work investigates when and how JEPAs can recover the underlying causal states from observations. We first introduce a latent variable model, in which high-dimensional observations are generated from latent causal states whose dynamics are governed by action-conditioned transition mechanisms. Based on this formulation, we develop a general information-theoretic objective that combines conditional likelihood maximization for learning transition dynamics with entropy maximization for preserving latent state information. We then establish identifiability conditions under which representations learned by this general objective recover the underlying latent causal states up to component-wise invertible transformations and permutation. One key condition for such identifiability is sufficient action-induced variation in the transition mechanisms. Guided by this finding, we instantiate the general objective with an action-modulated Gaussian additive-noise model, yielding action-modulated JEPA (A-JEPA). Experiments on synthetic environments verify the theoretical findings under the identifiability conditions and robustness to moderate violations, while visual benchmarks demonstrate improved state recovery and transfer to unseen transition mechanisms.

↑