arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SplitJEPA:无需重构学习不变与变异性潜在世界

SplitJEPA: Learning Invariant and Variant Latent Worlds without Reconstruction

Ruijin Hua, Zichuan Liu, Zhuokai Zhao, Yujia Zheng

arXiv 2610.12349首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Carnegie Mellon University; University of Chicago(伊利诺伊大学厄巴纳-香槟分校; 卡内基梅隆大学; 芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SplitJEPA是一种无需重构的联合嵌入预测架构,可直接在表示空间中联合恢复潜在状态的不变与变异组织,经实验验证其在鲁棒性和效率上具有实用价值。

AI 中文摘要

理解动态世界所需的远不止是总结其观测结果的潜在状态:该状态还应被组织为在相关观测之间保持共享的因子和在它们之间变化的因子。例如,当相机移动或光线变暗时,推动立方体到达目标的机器人应采取相同的动作,因为场景中没有任何物体移动。现有的这种分解方法通常通过重构来实现,因此潜在变量必须首先解释整个观测世界,然后其组织才能被信任。联合嵌入预测架构(JEPAs)直接对潜在状态建模,从不进行重构,但现有结果均未恢复其学习到的状态的不变和变异部分。因此,如何在不付出重构代价的情况下学习潜在世界的不变-变异结构仍然是一个开放问题。为了缩小这一差距,我们引入了SplitJEPA,这是一种直接在表示空间中联合恢复潜在状态及其不变和变异组织的JEPA,无需任何重构。我们证明,在平稳高斯预测动力学和满秩变异条件下,SplitJEPA可识别不变和变异子空间,直至独立的块式等距变换,且无需引入观测解码器。由于该保证不需要解码器,因此该结果将无重构的潜在恢复扩展到了不变-变异块识别。在合成非线性系统和机器人操纵任务上的实验支持了理论结果,并展示了其在鲁棒性和效率方面的实用价值。

英文摘要

Understanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that vary between them. For example, a robot pushing a cube to a goal should take the same action when the camera shifts or the lights dim, since nothing in the scene has moved. Existing approaches to this decomposition commonly obtain it through reconstruction, so the latent variables must first explain the entire observational world before their organization can be trusted. Joint embedding predictive architectures (JEPAs) model the latent state directly and never reconstruct, yet no existing result recovers the invariant and variant parts of the state they learn. How to learn the invariant-variant structure of the latent world without paying for its reconstruction therefore remains open. To close this gap, we introduce SplitJEPA, a JEPA that jointly recovers the latent state and its invariant and variant organization directly in representation space, without any reconstruction. We prove that, under stationary Gaussian predictive dynamics and a full-rank variation condition, SplitJEPA identifies the invariant and variant subspaces up to independent block-wise isometries, without introducing an observation decoder. Since the guarantee needs no decoder, the result extends reconstruction-free latent recovery to invariant-variant block identification. Experiments on synthetic nonlinear systems and robotic manipulation tasks support the theoretical results and show their practical value for both robustness and efficiency.

Comments23 pages, 15 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑