EVO-WAM:通过视频-动作验证进化世界动作模型
EVO-WAM: Evolving World Action Models through Video-Action Verification
- HITSZ(哈尔滨工业大学(深圳))
- SLAI
- THU(清华大学)
- HKUST(香港科技大学)
- PKU(北京大学)
- SJTU(上海交通大学)
- HKUSTGZ(香港科技大学(广州))
- HKU(香港大学)
- UBC(不列颠哥伦比亚大学)
- CUHK(香港中文大学)
- CUHKSZ(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EVO-WAM通过视频-动作验证,利用世界动作模型自身生成的轨迹进行迭代训练,无需外部执行,在未见任务上将成功率提升至2.5倍,并显著改善真实世界长时程任务表现。
AI中文摘要:
在不收集额外专家演示的情况下改进机器人策略以应对新任务,仍然是机器人学习中的一个核心挑战。世界动作模型(WAMs)利用广泛的视频先验来联合预测未来视频和动作,为适应新任务提供了一种潜在的监督来源。然而,生成的视频可能无法描绘任务完成情况,即使视觉上成功的视频也可能与不一致的动作配对,导致执行失败。我们提出EVO-WAM,一种通过从自身生成的视频-动作轨迹中学习来适应未见任务的框架,无需在外部环境中执行候选动作。首先,我们通过状态预测和锚定多帧上下文增强WAM训练,使其能够在没有外部执行反馈的情况下进行完整的自回归展开。其次,我们通过使用视觉语言模型选择任务完成前缀,并使用逆动力学模型验证其视频-动作一致性,来识别可靠的训练经验。第三,我们在验证的前缀上迭代训练WAM,并使用更新后的模型生成新的展开。在七个未见过的RoboTwin 2.0任务上,EVO-WAM将Cosmos3的平均成功率从26.9%提高到68.0%,将DreamZero的平均成功率从28.5%提高到46.4%,分别达到其初始成功率的约2.5倍和1.6倍。在真实世界的三个未见过的长时程复合任务上,它将Cosmos3的平均成功率从20.0%提高到76.7%,提升了56.7个百分点。项目页面:此https URL。
英文摘要:
Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.