arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LeWAM:一种基于扩散引导模型预测控制(MPC)的JEPA世界动作模型

LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC

Shashank Hegde, Alexander Popov, Elie Aljalbout, Nikolai Smolyanskiy

arXiv 2610.12407首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出LeWAM,一种基于无解码器JEPA隐空间的双向Transformer,其在状态读取、动作表现、规划能力上均优于常规WAM,兼具世界模型功能。

AI 中文摘要

世界动作模型(WAMs)可预测动作与未来观测结果,但其通常基于含噪声、冗余信息的重建式表示,会增加下游预测的复杂度。本文提出LeWAM,这是一种双向Transformer,可用于正向、反向动力学及策略预测,基于无解码器的JEPA隐空间,通过全部四种模式进行端到端训练。其优势如下:1)对齐性:线性探针从LeWAM的隐空间读取机器人与物体状态的效果优于常规Le世界模型(一种仅正向的JEPA世界模型),且该隐空间对视觉干扰项的忽略程度与LeWM相当,远优于重建式WAM;2)动作表现:LeWAM的闭环评估结果与在相同编码器、相同规模下训练的常规流匹配策略相当,同时还能提供世界模型;3)规划能力:WAM规划时采样原始动作会使MPC利用动力学模型的不准确性,而在策略头的噪声空间中规划可提升这类WAM的闭环性能。

英文摘要

World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at matched size, while also providing a world model. 3) Planning: Sampling raw actions when planning with WAMs lets MPC exploit dynamics-model inaccuracies; planning in the noise space of the policy head instead improves the closed-loop performance of these WAMs.

Comments14 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑