arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在JEPA世界模型中通过逆动力学保留不稳定模态

Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models

Leonardo F. Toso, Yann LeCun, James Anderson, Oumayma Bounou

arXiv 2610.07540首次发表:更新:

发表机构

Columbia University; New York University; AMI Labs(哥伦比亚大学; 纽约大学; AMI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对JEPA世界模型中不稳定模态坍缩导致控制失效的问题,提出逆动力学损失增强控制感知表示,理论证明其保留可达子空间,并在非线性视觉控制任务中验证有效性。

AI 中文摘要

机器人系统通常表现出不稳定模态,沿这些模态的小扰动和干扰可能导致无界增长,除非通过反馈进行校正。从高维视觉观测中控制此类系统需要保留这些模态的表示。联合嵌入预测架构(JEPAs)为从视觉数据学习此类表示及其动力学提供了自然框架。然而,我们证明,下一步预测结合防坍缩正则化并不能保证可控的不稳定模态被保留:训练损失可以在这些模态坍缩时被最小化,使得从学习到的表示进行稳定化变得不可能。为解决此问题,我们用动作重建目标(即逆动力学损失)增强世界模型训练,该目标鼓励控制感知表示,即保留控制关键特征的视觉表示。我们证明精确的动作重建使编码器在有限时域可达子空间上是单射的。因此,编码器不能丢弃任何在$H$步内由动作序列可达的状态方向。此外,我们表明,随着$H$增长,有限时域可控性格拉姆矩阵的主特征空间收敛到可控不稳定子空间。我们为线性系统建立了理论结果,并实证表明我们的发现扩展到非线性视觉控制任务(CartPole、Walker2D和PointMaze),突显了控制感知表示学习的优势。

英文摘要

Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback. Controlling such systems from high-dimensional visual observations requires representations that preserve these modes. Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning such representations and their dynamics from visual data. However, we demonstrate that next step prediction combined with anti-collapse regularization does not guarantee that controllable unstable modes are preserved: the training loss can be minimized while these modes are collapsed, making stabilization from the learned representation impossible. To address this, we augment world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control. We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace. Thus, the encoder cannot discard any state direction reachable by an action sequence within $H$ steps. Moreover, we show that, as $H$ grows, the dominant eigenspace of the finite-horizon controllability Gramian converges to the controllable unstable subspace. We establish our theoretical results for linear systems and demonstrate empirically that our findings extend to nonlinear visual control tasks (CartPole, Walker2D, and PointMaze), highlighting the benefits of control-aware representation learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑