发表机构
University of Illinois Urbana-Champaign; Cold Spring Harbor Laboratory; Brown University; Max Planck Institute for Intelligent Systems; ELLIS Institute Tübingen(伊利诺伊大学厄巴纳-香槟分校; 冷泉港实验室; 布朗大学; 马克斯·普朗克智能系统研究所; ELLIS图宾根研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DSReg利用结构多样性,通过依赖稀疏正则化,在无需重建、解码器或标签的情况下,可证明地恢复个体世界潜变量,并建立了首个完全可识别的JEPA,提升下游性能。
AI 中文摘要
从非线性ICA到字典学习和因果表示学习,恢复世界个体潜变量的方法通过重建、辅助监督或非高斯性等分布不对称性将潜变量锚定到观测上。没有这些锚定的方法,包括联合嵌入预测架构(JEPAs),只能将潜在状态识别到线性变换的程度,因此个体潜变量仍然混合。我们弥合了这一差距:个体世界潜变量可以在没有重建、没有解码器、没有标签的情况下被可证明地恢复。关键条件是结构多样性:不同的潜变量在观测上留下不同的依赖足迹,正如没有两片雪花是相同的。基于LeJEPA提供的线性可识别性,我们证明在结构多样性下,DSReg(依赖稀疏正则化)可以恢复个体世界潜变量,达到符号置换的程度,无需重建或解码器。它适用于任何线性可识别的表示的事后处理,重用训练检查点,且相对于联合训练没有损失,并建立了第一个完全可识别的JEPA,能够恢复每一个世界潜变量。此外,作为依赖足迹的条件,结构多样性严格弱于先前可识别潜变量模型的所有结构条件。在合成机制、世界模型探针、学习视觉编码器和外部渲染器中,DSReg保持了密集预测,同时改善了个体潜变量恢复和下游使用,并具有规模性。
英文摘要
Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.
CommentsProject page: https://dsreg.github.io/