人类作为人形机器人:通过人类对齐的具身从自我-外部人类视频实现零样本人形机器人学习
Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- DeepCybo
- ZGCA
- ZGCI
- Harbin Institute of Technology(哈尔滨工业大学)
- Huazhong University of Science and Technology(华中科技大学)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出Human-as-Humanoid框架,通过对齐机器人本体、感知设置和动作标签接口,将大规模人类演示视频转化为可执行的动作监督,实现高自由度人形机器人的零样本学习。
AI中文摘要:
跨机器人本体的视觉-语言-动作(VLA)模型需要高质量的观察-动作监督来学习可部署的动作分布,但扩展此类机器人数据仍然困难,尤其是对于高自由度的人形机器人。遥操作提供控制器对齐的监督,而人类自我中心视频捕捉多样的双手操作,但并未直接提供可执行的机器人动作。我们引入Human-as-Humanoid,一个人到人形机器人的监督框架,能够实现近实时的人类中心动作生成,通过联合对齐机器人本体、感知设置和动作标签接口,使人类演示可用于高自由度人形机器人VLA训练。基于PrimeU(一个人类对齐的60自由度上半身人形机器人),Human-as-Humanoid使用同步的自我-外部视频将部署对齐的自我中心观察与外部运动恢复配对,通过分阶段逆运动学(IK)将恢复的人类运动重定向为控制器对齐的60自由度动作块,并使用前向运动学(FK)感知的监督训练VLA模型,以保留手腕和指尖的任务空间几何。这将大规模人类演示从视觉观察转换为目标人形机器人的可执行观察-动作监督。实验在运动恢复、机器人动作空间和真实机器人部署层面验证了转换链。Human-as-Humanoid在我们的数据收集分析中比人形机器人遥操作实现了4.8-7.2倍的原始演示吞吐量增益,并且在多个下游任务中,仅使用转换后的人类标签进行后训练的策略能够泛化到真实机器人部署,无需目标任务机器人演示。官方项目网站见此https URL。
英文摘要:
Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot data remains difficult, especially for high-DoF humanoids. Teleoperation provides controller-aligned supervision, while human egocentric videos capture diverse bimanual manipulation but do not directly provide executable robot actions. We introduce Human-as-Humanoid, a human-to-humanoid supervision framework that enables near-real-time human-centric action generation, making human demonstrations usable for high-DoF humanoid VLA training by jointly aligning the robot embodiment, the sensing setup, and the action-label interface. Built on PrimeU, a human-aligned 60-DoF upper-body humanoid, Human-as-Humanoid uses synchronized ego-exo videos to pair deployment-aligned egocentric observations with exocentric motion recovery, retargets the recovered human motion through staged Inverse Kinematics (IK) into controller-aligned 60-DoF action chunks, and trains the VLA model with Forward Kinematics (FK)-aware supervision to preserve wrist and fingertip task-space geometry. This converts large-scale human demonstrations from visual observations into executable observation--action supervision for the target humanoid. Experiments validate the conversion chain at the motion-recovery, robot-action-space, and real-robot deployment levels. Human-as-Humanoid yields a 4.8--7.2x raw demonstration-throughput gain over humanoid teleoperation in our data-collection analysis, and on several downstream tasks, policies post-trained only with the converted human labels generalize to real-robot deployment without target-task robot demonstrations. The official project website is available at https://zgc-embodyai.github.io/Human-as-Humanoid.