arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

$R^2$-WAM:世界动作模型的修复与拒绝后训练

$R^2$-WAM: Repair-and-Reject Post-Training for World Action Models

Ruiyan Xu, Haisheng Su, Sixu Lin, Zhaokun Yue, Chengming Hu, Xin Jin, Guiliang Liu

arXiv 2610.04913首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; Manifold AI; Southeast University(香港中文大学(深圳); Manifold AI; 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出$R^2$-WAM两阶段后训练框架,通过修复预测与动作一致性并拒绝劣质样本进行负微调,提升世界动作模型的策略优化能力,在RoboTwin 2.0和真实折叠衬衫任务上显著优于基线。

AI 中文摘要

世界动作模型(WAMs)作为一种有前景的基础模型,通过预测采样动作的后果来支持策略优化。然而,如果视觉上看似合理的预测未能反映输入动作,则可能误导策略优化。为解决这一不匹配问题,我们提出了$R^2$-WAM,一个两阶段的修复与拒绝后训练框架,首先提高预测未来与输入动作的一致性,然后利用这些未来选择较差的动作样本进行负微调。修复阶段通过运动学对齐分数将想象扎根于观察到的机器人行为,该分数衡量预测运动与演示运动之间的一致性,使预测视频能够忠实反映其输入动作。使用修复后的视频模型,拒绝阶段比较采样动作与演示动作的想象结果,选择性地对预测任务进展低于演示参考值规定裕度的样本进行负微调。两个阶段共同将视频预测从表示学习扩展到基于后果的策略优化,无需额外的环境交互或改变推理过程。$R^2$-WAM在RoboTwin 2.0的干净和随机设置下平均成功率达到93.8%。在长时程真实世界折叠衬衫任务中,其平均成功率达到87.5%,而Fast-WAM为0%。

英文摘要

World Action Models (WAMs) emerge as a promising foundation for policy refinement by predicting the consequences of sampled actions. However, visually plausible predictions can mislead policy refinement if they fail to reflect the input actions. To address this mismatch, we introduce $R^2$-WAM, a two-stage repair-and-reject post-training framework that first improves the consistency of predicted futures with input actions, then uses these futures to select inferior action samples for negative fine-tuning. The repair stage grounds imagination in observed robot behavior through a kinematic alignment score that measures agreement between predicted and demonstrated motion, enabling the predicted video to faithfully reflect its input actions. Using the repaired video model, the rejection stage compares imagined outcomes of sampled and demonstrated actions, selectively applying negative fine-tuning to samples whose predicted task progress falls below the demonstrated reference by a prescribed margin. Together, the two stages extend video prediction from representation learning to consequence-based policy refinement without additional environment interaction or changes to the inference procedure. $R^2$-WAM achieves 93.8% average success on RoboTwin 2.0 across clean and randomized settings. On the long-horizon real-world Fold Shirt task, it achieves 87.5% average success, compared with 0% for Fast-WAM. Our project page is available at https://r2-wam.github.io/.

Comments24 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑