潜在演化世界动作模型
Latent evolving World Action Model
浏览论文内容
中文总结 AI 辅助
提出LeWAM,利用JEPA嵌入替代视频扩散骨干进行动作生成与环境演化预测,并引入DemoDPO离线偏好优化,以0.4B参数在RoboTwin 2.0上取得92.28%成功率。
中文摘要 AI 辅助
世界动作模型(WAMs)联合建模动作生成与环境动态,且大多基于预训练的视频扩散模型(VDMs)构建。在基于VDM的WAMs中,观测首先由VAE编码,产生的压缩潜变量随后由大型视频扩散骨干网络处理,以提取用于动作生成的有效特征。然而,这种范式将WAM的性能和训练成本与大规模视频生成预训练绑定,限制了WAM的效率和可扩展性。在本文中,我们从理论和实证两方面研究了视觉表示如何影响WAM中的动作生成。我们的结果表明,来自联合嵌入预测架构(JEPA)编码器的预测性嵌入比压缩的VAE潜变量更能支持动作生成,其中I-JEPA在我们的编码器比较中表现最佳。基于这些发现,我们提出了LeWAM,它基于JEPA嵌入来条件化动作生成,并通过预测同一空间中的未来嵌入来建模环境演化,而不依赖视频扩散骨干网络。我们进一步发现,模仿学习能够匹配示范动作,但无法区分更好的动作与更差的动作,尽管小的动作偏差可能极大地影响任务成功率。为了解决这一局限,且无需额外的环境交互或用于重置和安全所需的人工监督,我们引入了示范引导的DPO(DemoDPO),这是一种离线偏好优化阶段,直接从示范中推导出偏好监督。仅用0.4B可训练参数,LeWAM在RoboTwin 2.0上达到了92.28%的平均成功率,与最先进的VLAs和WAMs相当,并在真实世界操作任务中保持了实际有效性。
英文摘要
World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effective features for action generation. However, this paradigm ties WAM performance and training cost to large-scale video generation pretraining, limiting WAM efficiency and scalability. In this paper, we theoretically and empirically investigate how visual representations affect action generation in WAMs. Our results show that predictive embeddings from Joint-Embedding Predictive Architecture (JEPA) encoders better support action generation than compressed VAE latents, with I-JEPA performing best in our encoder comparison. Based on these findings, we propose LeWAM, which conditions action generation on JEPA embeddings and models environment evolution by predicting future embeddings in the same space, without relying on a video diffusion backbone. We further find that imitation learning matches demonstrated actions but does not distinguish better actions from worse ones, even though small action deviations can greatly affect task success. To address this limitation without additional environment interaction or the human oversight required for resets and safety, we introduce Demonstration-Guided DPO (DemoDPO), an offline preference refinement stage that derives preference supervision directly from demonstrations. With only 0.4B trainable parameters, LeWAM achieves an average success rate of 92.28\% on RoboTwin 2.0, comparable to that of state-of-the-art VLAs and WAMs, and maintains practical effectiveness on real-world manipulation tasks.
发表机构
- Zhejiang University(浙江大学)
- Westlake University(西湖大学)
- Baidu Inc.(百度公司)
机构由 AI 辅助整理,请以论文原文为准。