arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20709cs.RO

MoWAM:面向高效世界动作模型的显式未来运动预测

MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

Jiayu Wang, Bin Zhu, Yue Yu, Jingjing Chen

首次发表
浏览论文内容

中文总结 AI 辅助

MoWAM用显式未来运动预测替代未来视频生成,通过混合Transformer联合预测运动与动作,实现高效推理时扩展,在LIBERO及真实任务中提升分布外鲁棒性和成功率。

中文摘要 AI 辅助

世界动作模型(WAMs)通过纳入未来动态来改进机器人策略学习,然而在推理时显式生成未来视频会引入大量计算开销。移除未来生成虽能提升效率,但会使未来动态仅隐式编码于观测特征中,从而在分布偏移下可能限制鲁棒性。我们提出MoWAM,一种高效的世界动作模型,用显式未来运动预测替代未来视频生成。MoWAM并非重建完整的未来场景,而是将结构化机器人运动建模为未来的紧凑抽象,捕捉在当前场景和交互约束下机器人预期如何演变。一种混合Transformer架构在训练期间学习未来视觉动态,同时联合预测运动与动作,从而在推理时完全移除视频生成,同时保留对未来的显式表示。紧凑的运动表示还通过采样多组运动与动作候选对,并利用运动感知的任务进度验证器进行选择,实现了高效的推理时扩展。在LIBERO、LIBERO-Plus以及真实世界操作任务上的实验表明,MoWAM取得了强劲的分布内性能、改进的分布外鲁棒性,以及比代表性WAM基线更高的平均真实世界成功率。此外,随着探索更多候选,性能进一步提升,证明显式未来运动为推理时扩展提供了有效且高效的基础。

英文摘要

World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.

发表机构

  • Fudan University(复旦大学)
  • Singapore Management University(新加坡管理大学)
  • Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑