具有关节状态-动作生成功能的人形机器人世界动作模型
Humanoid World Action Model With Joint State--Action Generation
浏览论文内容
中文总结 AI 辅助
针对人形机器人系统的动作-执行差距问题,提出HWAM模型,通过联合生成参考动作与执行后身体状态,在LimX OLI人形机器人的糖果采摘等任务中取得优于Fast-WAM的成功率。
中文摘要 AI 辅助
人形机器人是通用操作的理想平台。近期的视觉-语言-动作(VLA)策略直接从多模态观测中学习动作,而世界动作模型(WAM)进一步融入未来视觉预测以改进动作生成。然而,在分层人形机器人系统中,VLA和WAM策略输出的参考动作随后通过全身控制、机器人动力学、平衡及接触来实现。这种分层结构造成了动作-执行差距:策略生成的参考动作与机器人实际执行的运动可能存在差异。若不显式建模实际执行的身体状态,未来的视觉预测必须同时解释场景演变以及参考动作与执行运动之间的差异,这使得将动作与其物理结果关联起来变得困难。我们提出HWAM,即具有关节状态-动作生成功能的人形机器人世界动作模型,该模型将机器人执行后的本体感受状态作为显式预测目标。通过联合生成参考动作及其实际执行的身体状态,HWAM将执行运动的监督直接融入动作学习。HWAM通过三条互补的条件路径进行训练:策略路径仅以当前观测为条件,联合对状态-动作轨迹进行去噪,以匹配部署条件;前向动力学建模(FDM)以动作和执行后状态为条件预测未来视觉观测;反向动力学建模(IDM)则从视觉转换中重构关节轨迹。这些路径共同连接策略参考、实际执行的身体运动和视觉结果。在LimX OLI人形机器人的三项真实机器人任务中,HWAM在评估的基线中达到最高成功率:在糖果采摘任务中,HWAM的成功率为70.6%,而Fast-WAM的成功率为43.3%。
英文摘要
Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action (VLA) policies learn actions directly from multimodal observations, while World Action Models (WAMs) further incorporate future visual prediction to improve action generation. However, in hierarchical humanoid systems, VLA and WAM policies output reference actions that are subsequently realized through whole-body control, robot dynamics, balance, and contact. This hierarchy creates an action--execution gap: the reference produced by the policy can differ from the motion realized by the robot. Without explicitly modeling the realized body state, future visual prediction must jointly explain scene evolution and discrepancies between reference actions and executed motion, making it difficult to associate an action with its physical outcome. We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. By jointly generating reference actions and their realized body states, HWAM directly incorporates supervision of executed motion into action learning. HWAM is trained through three complementary conditional paths. The Policy path jointly denoises state--action trajectories conditioned only on current observations, matching deployment conditions. Forward Dynamics Modeling (FDM) predicts future visual observations conditioned on actions and post-execution states, while Inverse Dynamics Modeling (IDM) reconstructs the joint trajectory from visual transitions. Together, these paths connect policy references, realized body motion, and visual outcomes. HWAM achieves the highest success rate among evaluated baselines on three real-robot tasks on the LimX OLI humanoid. On Candy Picking, HWAM achieves a 70.6% success rate, compared with 43.3% for Fast-WAM.
发表机构
- LimX Dynamics(灵蜥动力)
- Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。