发表机构
Across Physics; Beijing SFlare Robotics Technology Co., Ltd.(跨物理; 北京星焰机器人科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单目视频人体到机器人上半身动作重定向的几何保持框架,通过统一体-手重建与形态无关几何迁移,显著降低手部重投影误差,并成功在物理机器人上演示。
AI 中文摘要
单目RGB视频为上半身机器人动作提供了易获取的人体演示来源,然而视频驱动的人体到机器人迁移仍具挑战性,因为身体和手部动作在不同空间尺度下恢复,人体与机器人运动学差异显著,且精细的远端动作难以跨实体保持。我们提出一种几何保持的动作重定向框架,该框架将统一的体-手重建与形态无关的几何迁移相结合。逐帧身体估计、视频级观测和详细的手部证据共同约束一个单一的可微动量人体骨架(MHR)状态,而瞬态手部伪影在参数空间中被修复。重建的动作由手臂段方向、肘部构型、相对手掌方向和双侧手腕关系表示,并通过多阶段逆运动学和机器人特定的手部适配在目标机器人上实现。在更广泛的系统中,Across-VAM提供视频生成,而Across-WAM执行人体到机器人的动作映射。该方法在16个单目视频(包含1,769个源帧)上评估,包括10个手语和6个伸手抓取序列。与SAM 3D Body相比,统一重建将平均手部重投影误差从22.36像素降至7.21像素。所有16个重定向轨迹均完成运动学仿真回放,代表性的手语和伸手抓取动作进一步在物理机器人上演示。结果表明,从单目人体视频到双臂灵巧机器人协调上半身动作的统一流程。
英文摘要
Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer. Frame-wise body estimates, video-level observations, and detailed hand evidence jointly constrain a single differentiable Momentum Human Rig (MHR) state, while transient hand artifacts are repaired in parameter space. The reconstructed motion is represented by arm-segment directions, elbow configuration, relative palm orientation, and bilateral wrist relations, and is realized on the target robot through multi-stage inverse kinematics and robot-specific hand adaptation. Within the broader system, Across-VAM provides video generation, whereas Across-WAM performs human-to-robot motion mapping. The method is evaluated on 16 monocular videos comprising 1,769 source frames, including 10 signing and six reach-to-grasp sequences. Unified reconstruction reduces mean hand reprojection error from 22.36 to 7.21 pixels relative to SAM 3D Body. All 16 retargeted trajectories completed kinematic simulation playback, and representative signing and reach-to-grasp motions were further demonstrated on a physical robot. The results demonstrate a unified pipeline from monocular human video to coordinated upper-body motion on a dual-arm dexterous robot.
Comments9 pages, 4 figures, 3 tables