AI 中文总结
研究受控世界模型的可识别性问题,建立在状态依赖高斯行为策略下的联合可识别性理论,识别出两个条件,证明满足条件时JEPA目标的全局极小值可识别潜在状态和受控转移,推导定量界限并通过实验验证理论。
AI 中文摘要
学习从高维观测中推断环境动态并在候选动作下预测结果的世界模型对规划和控制至关重要。联合嵌入预测架构(JEPAs)为在表示空间中学习此类模型提供了一个引人注目的框架。最近的动作条件扩展在视觉控制和潜在空间规划中表现出良好前景,但一个基本问题未解决:何时受控潜在预测能识别潜在状态和受控动态?在非线性观测和条件动作变化有限的行为策略下这具有挑战性。我们建立了在状态依赖高斯行为策略下具有高斯潜在状态的受控世界模型的联合可识别性理论。我们识别出两个依赖策略的条件:可预测信号的谱分离控制表示可识别性,而非退化条件动作变化控制转移可识别性。我们证明当两个条件都成立时,JEPA目标的每个全局极小值能识别潜在状态和受控转移直至正交变换。我们还推导了近似优化下表示和转移可识别性的定量界限。最后,我们沿着弱激发动作方向构造预测器扰动,其反事实与实际策略误差比是转移可识别性余量的倒数,揭示了有限动作覆盖的代价。跨非线性观测图和行为策略的实验证实了该理论,并展示了对转移可识别性、反事实预测和目标条件潜在规划的影响。
英文摘要
World model serves as a promising tool to infer environment dynamics under high-dimensional observations and candidate actions. Recently, LeCun's JEPA provides a compelling framework for learning such models in representation space. Its action-conditioned extension plays a central role in visual control and latent-space planning, but leaves a fundamental question: can it recover the controlled dynamics from nonlinear observations? This paper presents a joint identifiability condition for controlled world models with Gaussian latent states, which consists of two coupled components: (1) representation identifiability and (2) transition identifiability. The former depends on the spectral separation property while the latter is related to non-degenerate variation of conditional action. We prove that when this condition holds, minimizing the LeJEPA-style predictive objective can recover both latent states and controlled dynamics in the sense of orthogonal transformation. We further prove that the upper bound of transition prediction error is inversely proportional to the spectral separation margin. We also characterize an attainable amplification of counterfactual prediction error that scales inversely with the weakest conditional action-excitation margin. The theoretical predictions are empirically supported across four nonlinear observation settings.