发表机构
The Chinese University of Hong Kong; The University of Hong Kong; Peking University(香港中文大学; 香港大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WholeBodyWAM通过联合预测视觉动态、操作动作和全身控制意图,将预训练的世界-动作先验泛化到人形机器人全身操作,在仿真和真实实验中显著提升了任务成功率与泛化能力。
AI 中文摘要
世界动作模型(WAMs)通过联合建模视觉动态和动作,为通用机器人操作提供了一种有前景的方法。然而,大多数WAM研究集中在桌面或以手臂为中心的操作上,而人形机器人的全身操作(loco-manipulation)仍较少被探索。为解决这一差距,我们提出了WholeBodyWAM,它联合预测未来的视觉动态、操作动作和全身控制意图,以实现可泛化的人形机器人全身操作。它保留了预训练的世界-动作先验,同时锚定异构全身控制器(WBC)语义并协调全身行为。大量实验表明,WholeBodyWAM在仿真任务中实现了91.9%的总体成功率,在真实世界分布外任务进展上提升了0.23,并且相对于各自的基线,在不同WBC间的成功率方差降低了70%。这些结果表明,通过结构化的WBC锚定和协调来扩展预训练的世界-动作先验,而非从头重新学习全身行为,为可扩展的人形机器人全身智能开辟了一条路径。项目页面:此https URL。
英文摘要
World Action Models (WAMs) offer a promising approach to general-purpose robot manipulation by jointly modeling visual dynamics and actions. However, most WAM studies focus on tabletop or arm-centric manipulation, while humanoid loco-manipulation remains less explored. To address this gap, we introduce WholeBodyWAM, which jointly predicts future visual dynamics, manipulation actions, and whole-body control intents for generalizable humanoid loco-manipulation. It preserves pre-trained world-action priors while grounding heterogeneous whole-body controller (WBC) semantics and coordinating whole-body behavior. Extensive experiments show that WholeBodyWAM achieves an overall simulation task success rate of 91.9%, with a 0.23 improvement in real-world out-of-distribution task progress and a 70% reduction in success-rate variance across WBCs relative to the respective baselines. These results suggest a path toward scalable humanoid whole-body intelligence by extending pre-trained world-action priors through structured WBC grounding and coordination, rather than relearning whole-body behavior from scratch. Project page: https://wholebodywam.github.io/.
Comments8 pages, 8 figures, 3 tables