arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34375cs.LGcs.AIcs.RO

LRC-JEPA:解耦动力学与残差上下文以构建高效世界模型

LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models

  • USC(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao

AI总结:

LRC-JEPA通过将动力学与残差上下文解耦,在紧凑潜在空间中实现高效规划,在四个模拟环境和真实Bridge-v2数据集上显著提升规划成功率并超越更大模型。

AI中文摘要:

紧凑的JEPA世界模型能够实现高效的潜在空间规划,但在无奖励自监督下训练的低维表示必须同时编码动作条件动力学和可预测的视觉上下文。这种竞争可能将可控状态与高秩的无关外观纠缠在一起,并在场景变得更复杂时降低规划性能。我们提出了LRC-JEPA,一种轻量级的端到端世界模型,它将信息路由到紧凑的预测性潜在变量$\mathbf{z}$和学习查询的残差上下文嵌入$\mathbf{u}$中。只有$\mathbf{z}$由动力学模型传播并用于规划,而$\mathbf{u}$捕获时间上持久的信息用于交叉注意力重建;一个可微分的残差连接鼓励潜在变量保留互补的动态内容。在明确的假设下,我们证明了所得表示是充分的、最小的、对无关因素不变的且解耦的。在四个模拟控制环境中,LRC-JEPA相比参数匹配的JEPA基线,平均规划成功率提高了9个百分点,并且匹配或超过了规模大得多的预训练模型。在真实世界的Bridge-v2数据集上,其5.5M参数的主动编码器优于DINO-WM(22.1M)和V-JEPA2(303.9M)编码器,同时实现更快的规划。物理状态探针、重建干预和消融实验证实了LRC-JEPA表示解耦的有效性。

英文摘要:

Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent $\mathbf{z}$ and learned-query residual-context embeddings $\mathbf{u}$. Only $\mathbf{z}$ is propagated by the dynamics model and used for planning, while $\mathbf{u}$ captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.

↑