发表机构
NYU; Duke University; Mila; Brown University(纽约大学; 杜克大学; 米拉; 布朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LpWM模型,以稀疏表示替代密集表示建模动作条件潜在动力学,在PushT任务上规划成功率优于密集模型,且能揭示可解释的动力学结构。
AI 中文摘要
联合嵌入预测架构(JEPAs)学习用于规划的潜在动力学,并通过将特征匹配到各向同性高斯等最大熵分布来避免表示崩溃,从而产生密集表示。然而,目前尚不清楚密集表示是否是建模动力学的最有利几何结构。在这项工作中,我们探究了不同的几何结构——稀疏表示是否能使动作条件潜在动力学更易于建模,以及从这类表示中会产生何种动力学结构。我们首先证明,在足够高维的独热潜在空间中,动作条件线性动力学可以任意好地近似非线性Lipschitz动力学,且随维度增长,滚动误差会消失。这促使分布式稀疏表示成为独热稀疏性的实用松弛。我们引入Lp世界模型(LpWM),这是一种使用修正分布匹配正则化(RDMReg)进行正则化的JEPA模型,用于将编码器特征匹配到修正广义高斯分布,从而产生非负稀疏编码。从经验上看,稀疏性降低了成功规划所需的预测器复杂度:在PushT任务上,在中等预测器容量下,稀疏LpWM的规划成功率比密集LeWM高出多达57%。这一优势也超出了高斯分布匹配的范围,LpWM在多个预测器族上的表现优于密集VICReg表示。我们进一步发现,所学的稀疏表示是模式分解的,其支持集编码离散动力学机制,而特征幅度捕获机制内的连续状态。综合来看,这些结果表明,稀疏表示可以降低控制所需的预测器复杂度,同时揭示可解释的结构。
英文摘要
Joint-embedding predictive architectures (JEPAs) learn latent dynamics for planning and avoid representation collapse by matching features to maximum-entropy distributions such as isotropic Gaussians, yielding dense representations. However, it is unclear whether dense representations are the most favorable geometry for modeling dynamics. In this work, we ask whether a different geometry, sparse representations, can make action-conditioned latent dynamics easier to model, and what dynamical structure emerges from such representations. We first show that nonlinear Lipschitz dynamics can be approximated arbitrarily well by action-conditioned linear dynamics in a sufficiently high-dimensional one-hot latent space, with rollout error vanishing as the dimension grows. This motivates distributed sparse representations as a practical relaxation of one-hot sparsity. We introduce LpWorldModel (LpWM), a JEPA model regularized with Rectified Distribution Matching Regularization (RDMReg) to match encoder features to a Rectified Generalized Gaussian distribution, yielding non-negative sparse codes. Empirically, sparsity lowers the predictor complexity required for successful planning: on PushT, sparse LpWM outperforms dense LeWM by up to 57% in planning success at intermediate predictor capacities. This advantage also extends beyond Gaussian distribution matching, with LpWM outperforming dense VICReg representations across multiple predictor families. We further find that the learned sparse representations are mode-factored, with support encoding discrete dynamical regimes and feature magnitudes capturing continuous within-regime state. Together, these results suggest that sparse representations can reduce the predictor complexity required for control while revealing interpretable structure.