arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21514cs.RO

Skel-WAM:一种手部骨骼条件化的世界动作模型,用于人机操作迁移

Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer

  • The Chinese University of Hong Kong(香港中文大学)
  • Shanghai AI Laboratory(上海人工智能实验室)
  • Joy Future Academy(京东探索研究院)
  • CPII under InnoHK(香港InnoHK下的CPII)

机构由 AI 辅助整理,请以论文原文为准。

Zetao Cai, Yaping Li, Yiqun Wang, Xinyu Zhan, Yuyin Yang, Haoxiang Ma, Kailin Li, Tao Lu, Jiangmiao Pang, Linning Xu, Dahua Lin

AI总结:

针对机器人演示成本高且覆盖有限的问题,Skel-WAM通过统一手部骨骼接口对齐人机运动,利用混合Transformer联合学习视觉与骨骼动态,在真实和模拟任务中显著提升成功率,并扩展任务覆盖。

AI中文摘要:

机器人演示数据的采集成本高昂,且通常对任务变化的分布覆盖有限。人类视频提供了一种低成本的补充操作经验来源,但从中学习需要弥合视觉外观和动作空间中的具身差异。我们提出了Skel-WAM,一种世界动作模型,通过统一的手部骨骼运动接口来弥合这些差异。关键洞察在于通过共同的手部拓扑结构对齐人与机器人的运动,将骨骼叠加(将运动锚定在场景中)与编码显式手部运动学的结构化2.5维关键点相结合。视频专家和关键点专家通过Transformer混合架构联合学习视觉和骨骼动态,而一个单独经机器人训练的Action Expert则将这些预测映射为可执行的控制指令。这种分离使得人类和机器人的演示能够直接监督共享的动态学习,而无需为人类视频提供机器人动作标签。在四个真实世界双臂任务和七个模拟任务中,Skel-WAM分别实现了79.86%和63.29%的平均成功率,分别比最强基线高出22.22和8.28个百分点。人机协同训练在机器人训练数据中未包含的任务变体上,将真实世界成功率从38.89%提升至86.11%,提高了一倍以上。这些结果表明,共享的骨骼接口能够实现人类和机器人数据的联合学习,并通过互补的人类演示扩展机器人的任务覆盖范围。

英文摘要:

Robot demonstrations are expensive to collect and often provide limited distributional coverage of task variations. Human videos offer a low-cost source of complementary manipulation experience, but learning from them requires bridging embodiment gaps in visual appearance and action spaces. We introduce Skel-WAM, a world action model that bridges these differences through a unified hand-skeleton motion interface. The key insight is to align human and robot motion through a common hand topology, combining skeleton overlays that ground motion in the scene with structured 2.5-D keypoints that encode explicit hand kinematics. Video and Keypoint Experts jointly learn visual and skeletal dynamics through a Mixture-of-Transformers, while a separate robot-trained Action Expert maps these predictions to executable controls. This separation enables human and robot demonstrations to directly supervise shared dynamics without requiring robot action labels for human videos. Across four real-world bimanual tasks and seven simulated tasks, Skel-WAM achieves average success rates of 79.86% and 63.29%, surpassing the strongest baseline by 22.22 and 8.28 percentage points, respectively. Human-robot cotraining more than doubles real-world success on task variations absent from robot training data, from 38.89% to 86.11%. These results demonstrate that a shared skeletal interface enables joint learning across human and robot data and expands robot task coverage through complementary human demonstrations.

↑