发表机构
Intelligent Space Robotics Lab, Skolkovo Institute of Science and Technology; R&D Center, MWS(智能空间机器人实验室,斯科尔科沃科学技术研究院; 研发中心,MWS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在解决人类以自我为中心视频用于类人机器人学习的难题,提出AgenticFocus混合现实合成流程,恢复物体几何形状、重建手部运动并重新定位,生成数据集,相比跨实体基线降低轨迹误差、使手腕运动更平滑。
AI 中文摘要
人类以自我为中心的视频是类人机器人策略学习的可扩展监督源,但当前的流程在手部与物体遮挡、过度简化的动作或专门的捕获硬件方面存在困难。我们引入了AgenticFocus,这是一种混合现实合成流程,通过恢复被遮挡物体的几何形状、重建全手运动,并通过相机相对对齐和分层合成将其重新定位到类人机器人实体,将普通的第一人称视角人类视频转换为机器人可训练的演示。生成的数据集将聚焦的视觉观察与同步的机器人动作和状态配对。与跨实体基线相比,AgenticFocus实现了更低的轨迹误差和更平滑的手腕运动,SPARC分数分别为-5.18,而基线为-5.56和-6.05。
英文摘要
Human egocentric video is a scalable supervision source for humanoid policy learning, but current pipelines struggle with hand-object occlusion, oversimplified motion, or specialized capture hardware. We introduce AgenticFocus, a Mixed Reality synthesis pipeline that converts ordinary first-person-view human videos into robot-trainable demonstrations by restoring occluded object geometry, reconstructing full-hand motion, and retargeting it to a humanoid embodiment through camera-relative alignment and layered compositing. The resulting dataset pairs focused visual observations with synchronized robot actions and states. AgenticFocus achieves lower trajectory error and smoother wrist motion than cross-embodiment baselines, with SPARC scores of -5.18 versus -5.56 and -6.05.