arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22364cs.AIcs.RO

WAM-OPD:面向世界动作模型的策略内蒸馏

WAM-OPD: Joint Video-Action Supervision for World Action Model Post-Training with On-Policy Distillation

Liuhaichen Yang, Zhengyang Zhong, Hanshang Zhu, Ningwei Bai, Qichen Yin, Zhi Han, Jiarui Qin, Zhuang Jiang, Chenchao Sheng, Hanbo Ma, Junkai Liu, Junkai Sun, Do… 展开作者

Liuhaichen Yang, Zhengyang Zhong, Hanshang Zhu, Ningwei Bai, Qichen Yin, Zhi Han, Jiarui Qin, Zhuang Jiang, Chenchao Sheng, Hanbo Ma, Junkai Liu, Junkai Sun, Dongcheng Lyu, Bo Liu, Yi Dong, Zezhi Tang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出WAM-OPD策略内蒸馏方案,用于修复世界动作模型学生模型,在RoboTwin 2.0两项机器人任务中提升了Flash-WAM的任务成功率,为视频优先WAM提供了有前景的后训练方法。

中文摘要 AI 辅助

世界动作模型(WAMs)将视觉未来预测与机器人动作生成相结合,但加速的学生模型在蒸馏过程中可能会丢失任务能力,后续还会遇到离线数据未充分表征的状态。本文研究策略内蒸馏(OPD)是否能在无需稀疏奖励强化学习的情况下修复此类学生模型。我们提出WAM-OPD,这是一种针对视频优先WAM的部署一致后训练方案:学生模型在环境中行动,因此决定了历史分布;冻结的教师模型为这些学生历史标记一致的视频和动作目标,同时学生动作分支在部署时的自身生成视频计划下进行训练;联合视频和动作损失会更新共享主干中的轻量适配器,同时结合动作流匹配正则化器。在RoboTwin 2.0的两项任务初步研究中,发布的单视频/单动作步Flash-WAM在HANDOVER MIC任务上的成功率从0.0%提升至58.3%,在PUT OBJECT CABINET任务上从16.7%提升至33.3%。这些特定任务结果是初步能力验证,而非广泛或均匀泛化的证据,但仍表明对学生诱导的历史进行密集教师监督是视频优先WAM的有前景后训练接口。

英文摘要

World Action Models (WAMs) generate both future video and robot actions, offering two connected outputs for post-training supervision. How can a pretrained WAM learn from a stronger Teacher on the histories it encounters during execution? We present WAM-OPD, which collects Student rollout histories and queries a Teacher for paired video and action targets. The Student learns from both targets while retaining its one-step video and action generation at deployment. Across 12 RoboTwin 2.0 tasks, WAM-OPD improves average success from 33.8% to 65.7%; across four real-robot tasks, it improves average success from 51.4% to 64.6%. With the collected Student histories held fixed, joint video-action supervision achieves the highest observed success on all three ablation tasks, while either modality alone also improves performance. A separate comparison with Teacher-generated histories finds task-dependent differences between the two history sources. These results demonstrate the value of paired video-action supervision for improving WAM policies without increasing their deployed sampling budget.

发表机构

  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑