发表机构
Harbin Institute of Technology; Dalian University of Technology; Southern University of Science and Technology(哈尔滨工业大学; 大连理工大学; 南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DEWO,一种部署后学习范式,通过视觉经验细化世界模型,结合动作模仿提升灵巧操作成功率,实验显示两轮学习后成功率提升20.7个百分点。
AI 中文摘要
世界-动作模型(World-Action Models, WAMs)将动作生成与物理交互如何展开的预测相结合。然而,当前部署后学习范式通常在不要求更好世界预测的情况下改善行为。特别是在灵巧操作中,小执行误差可能在高维动作空间中累积,阻碍策略改进并将交互推离世界模型的训练分布。受此启发,我们提出直接经验世界模型优化(Direct Experience World-Model Optimization, DEWO),一种针对WAMs的部署后学习范式,它在动作模仿的同时,通过视觉经验细化世界表征,以更好地调节动作生成。具体而言,它识别交互转折点,并从成功和失败的未来中学习以支持无分类器引导。一个额外的价值头从视频表征估计任务进展,并在推理期间进展停滞时激活引导。在五个DexJoCo任务中,DEWO在所有三种WAM公式上提高了平均成功率。消融实验表明,来自成功和失败延续的视觉监督在预测和控制方面均优于仅动作监督。在Wuji和Sharpa上的四个真实世界任务中,3×3网格评估显示,两轮部署学习将至少有一次初始成功的单元格中的成功率从51.0%提高到71.7%,增益为20.7个百分点。这些发现支持通过部署经验持续进行预测学习以改善控制,使世界建模成为WAM适应的主动部分。
英文摘要
World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.
CommentsWorld Action Model; post-deployment training; dexterous manipulation