arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

实现你所想象的:学习将动作与视觉计划对齐

Achieve What You Imagined: Learning to Align Actions with Visual Plans

Yuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo, Yichen Luo, Zhibo Pang, Peng Zhou, Sichao Liu

arXiv 2609.33832首次发表:更新:

发表机构

KTH; Beihang University; The Hong Kong University of Science and Technology (Guangzhou); Peking University; Great Bay University(瑞典皇家理工学院; 北京航空航天大学; 香港科技大学(广州); 北京大学; 大湾区大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对世界-动作模型视觉预测与动作后果不一致的问题,提出利用冻结动作条件世界模型构建反馈,通过流策略优化对齐动作与视觉计划,在UR5任务上将成功率从43.4%提升至75.1%。

AI 中文摘要

世界-动作模型可以联合预测未来的视觉观察和机器人动作。然而,其视觉预测与生成动作所隐含的后果之间可能存在差异。我们观察到,世界-动作模型在生成能够可靠实现任务完成结果的动作序列之前,往往能先产生视觉上看似合理的任务完成结果。因此,我们将世界-动作模型生成的视觉预测视为目标条件视觉提案,而非可直接执行的计划。我们使用冻结的动作条件世界模型来预测动作条件的后果,并基于两个未来预测之间的一致性以及与最终目标的对齐来构建反馈。利用这一反馈,我们采用流策略优化(Flow Policy Optimization, FPO)来优化世界-动作模型的动作头。该框架避免了在线机器人交互以及额外训练任务特定奖励模型的需求。在四个真实世界的UR5操作任务中,我们的方法将平均成功率从43.4%提高到75.1%,而π0.5的成功率为61.4%。这些结果表明,跨模型预测差异可以为改进所评估操作任务下的机器人策略提供有用的反馈。网站:此HTTPS URL

英文摘要

World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for $π_{0.5}$. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/

Comments9 pages, 8 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑