arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CF-JEPA:通过可控性分解提升JEPA世界模型的鲁棒性

CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization

Morgan Byrd, Robert Wright, Sehoon Ha

arXiv 2610.00727首次发表:更新:

发表机构

Georgia Tech Research Institute; Georgia Institute of Technology(佐治亚理工学院研究所; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CF-JEPA,通过将潜在空间分解为可控与不可控子空间,提升JEPA世界模型在视觉控制任务中对干扰的鲁棒性,避免潜在坍缩,并在2D/3D及机器人任务中验证了性能提升。

AI 中文摘要

控制具有视觉的智能体需要能够将有用信息与无关背景信息分离。JEPA风格的潜在世界模型似乎是实现这一目标的自然方法,因为它们不进行像素级重建;然而,它们仍然对这些干扰信号敏感,并经历潜在空间坍缩。在这项工作中,我们引入了可控性分解JEPA(CF-JEPA),这是一种JEPA风格的世界模型,它将潜在空间分解为可控和不可控子空间。这种分解使我们能够将所有干扰信息捕获到不可控区域,同时我们将与控制相关的潜在信息用于任务。通过这种方式,我们展示了在标称条件下2D和3D控制任务的性能相当,在受干扰条件下性能提升,其中CF-JEPA是唯一不经历潜在空间坍缩的模型。我们还在受干扰条件下对模拟机器人任务验证了我们的模型,突出了这种方案的实际应用。

英文摘要

Controlling an agent with vision requires being able to separate useful information from irrelevant background information. JEPA-style latent world models seem like a natural approach for this, as they do not perform pixel-level reconstruction; however, they are still sensitive to these distractor signals and experience latent collapse. In this work, we introduce Controllability Factorized JEPA (CF-JEPA), a JEPA-style world model which splits the latent space into controllable and uncontrollable subspaces. This factorization allows us to capture all the distractor information into the uncontrollable region, while we use the control-relevant latent information for our task. With this, we show comparable performance across 2D and 3D control tasks under nominal conditions and improved performance under distracted conditions, where CF-JEPA is the only model that does not experience latent collapse. We also validate our model under distracted conditions for a simulated robot task, highlighting the practical application of such a scheme.

CommentsWebsite: https://morganbyrd03.github.io/cf-jepa/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑