WAM-OPD:面向世界动作模型的策略内蒸馏
WAM-OPD: Joint Video-Action Supervision for World Action Model Post-Training with On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
本研究提出WAM-OPD策略内蒸馏方案,用于修复世界动作模型学生模型,在RoboTwin 2.0两项机器人任务中提升了Flash-WAM的任务成功率,为视频优先WAM提供了有前景的后训练方法。
中文摘要 AI 辅助
世界动作模型(WAMs)将视觉未来预测与机器人动作生成相结合,但加速的学生模型在蒸馏过程中可能会丢失任务能力,后续还会遇到离线数据未充分表征的状态。本文研究策略内蒸馏(OPD)是否能在无需稀疏奖励强化学习的情况下修复此类学生模型。我们提出WAM-OPD,这是一种针对视频优先WAM的部署一致后训练方案:学生模型在环境中行动,因此决定了历史分布;冻结的教师模型为这些学生历史标记一致的视频和动作目标,同时学生动作分支在部署时的自身生成视频计划下进行训练;联合视频和动作损失会更新共享主干中的轻量适配器,同时结合动作流匹配正则化器。在RoboTwin 2.0的两项任务初步研究中,发布的单视频/单动作步Flash-WAM在HANDOVER MIC任务上的成功率从0.0%提升至58.3%,在PUT OBJECT CABINET任务上从16.7%提升至33.3%。这些特定任务结果是初步能力验证,而非广泛或均匀泛化的证据,但仍表明对学生诱导的历史进行密集教师监督是视频优先WAM的有前景后训练接口。
英文摘要
World Action Models (WAMs) generate both future video and robot actions, offering two connected outputs for post-training supervision. How can a pretrained WAM learn from a stronger Teacher on the histories it encounters during execution? We present WAM-OPD, which collects Student rollout histories and queries a Teacher for paired video and action targets. The Student learns from both targets while retaining its one-step video and action generation at deployment. Across 12 RoboTwin 2.0 tasks, WAM-OPD improves average success from 33.8% to 65.7%; across four real-robot tasks, it improves average success from 51.4% to 64.6%. With the collected Student histories held fixed, joint video-action supervision achieves the highest observed success on all three ablation tasks, while either modality alone also improves performance. A separate comparison with Teacher-generated histories finds task-dependent differences between the two history sources. These results demonstrate the value of paired video-action supervision for improving WAM policies without increasing their deployed sampling budget.
发表机构
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。