发表机构
HKUST(GZ); Noitom Robotics; Hanyang University; HKUST; HKU(香港科技大学(广州); 诺亦腾机器人公司; 汉阳大学; 香港科技大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出UMR框架,通过学习密集点云对应关系实现无需人工映射的人形机器人统一运动重定向,在多场景下比现有最优方法实现更高运动保真度,为大规模人类运动数据转机器人训练数据提供可扩展基础。
AI 中文摘要
人形机器人学习越来越依赖于将海量多样的人类运动数据转化为高质量的机器人参考轨迹。然而,由于人类与机器人之间在形态、自由度、关节范围和运动学约束方面存在巨大差异,将人类运动重定向到人形机器人极具挑战性。现有的重定向方法通常通过手工设计的稀疏关键点或身体部位对来定义人与机器人的对应关系,因此重定向质量高度依赖于人工语义设计,限制了其在不同运动源和机器人形态间的可扩展性,且仅能为复现精细姿态和交互提供稀疏指导。本文提出了统一运动重定向(UMR)框架,该框架学习密集点云对应关系,无需人工设计的人与机器人映射。通过将外部点云视为人类运动与人形机器人之间的统一接口,UMR将重定向与特定源的骨骼语义及特定机器人的拓扑结构解耦。学习到的密集对应关系为约束点云匹配优化提供了细粒度几何锚点,实现了表面级姿态对齐和交互接触的直接传递。实验表明,UMR可统一异构运动源、机器人实体以及从移动到交互等下游场景的重定向,且比现有最优方法实现了更高的运动保真度和合理性。因此,UMR为将大规模人类运动参考转化为可用于机器人的训练数据提供了可扩展的基础。
英文摘要
Humanoid learning increasingly relies on transforming vast and diverse human motion data into high-quality robot reference trajectories. However, retargeting human motion to humanoid robots is challenging due to substantial differences in morphology, degrees of freedom, joint ranges, and kinematic constraints between humans and robots. Existing retargeting methods typically address these differences by defining human-robot correspondence through hand-crafted sparse keypoints or body-part pairs. As a result, retargeting quality depends heavily on manual semantic design, limiting scalability across motion sources and robot morphologies and providing only sparse guidance for reproducing detailed poses and interactions. In this paper, we present Unified Motion Retargeting (UMR), a framework that learns dense point cloud correspondence without requiring manually designed human-robot mappings. By treating exterior point clouds as a unified interface between human motion and humanoid robots, UMR decouples retargeting from source-specific skeletal semantics and robot-specific topology. The learned dense correspondence provides fine-grained geometric anchors for constrained point cloud matching optimization, enabling surface-level pose alignment and direct transfer of interaction contacts. Experiments demonstrate that UMR unifies retargeting across heterogeneous motion sources, robot embodiments, and downstream scenarios ranging from locomotion to interaction, while achieving higher motion fidelity and plausibility than state-of-the-art methods. UMR therefore provides a scalable foundation for transforming large-scale human motion references into robot-ready training data.