arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgenticFocus:从人类第一人称视角视频进行物体保留的混合现实合成,用于灵巧人形机器人学习

AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning

Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov, Jeffrin Sam, Dmitrii Iarchuk, Yara Mahmoud, Daniia Zinniatullina, Mikhail Konenkov, Dzmitry Tsetserukou

arXiv 2607.08857首次发表:更新:

发表机构

Intelligent Space Robotics Lab, Skolkovo Institute of Science and Technology; R&D Center, MWS(智能空间机器人实验室,斯科尔科沃科学技术研究院; 研发中心,MWS)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在解决人类以自我为中心视频用于类人机器人学习的难题,提出AgenticFocus混合现实合成流程,恢复物体几何形状、重建手部运动并重新定位,生成数据集,相比跨实体基线降低轨迹误差、使手腕运动更平滑。

AI 中文摘要

人类以自我为中心的视频是类人机器人策略学习的可扩展监督源,但当前的流程在手部与物体遮挡、过度简化的动作或专门的捕获硬件方面存在困难。我们引入了AgenticFocus,这是一种混合现实合成流程,通过恢复被遮挡物体的几何形状、重建全手运动,并通过相机相对对齐和分层合成将其重新定位到类人机器人实体,将普通的第一人称视角人类视频转换为机器人可训练的演示。生成的数据集将聚焦的视觉观察与同步的机器人动作和状态配对。与跨实体基线相比,AgenticFocus实现了更低的轨迹误差和更平滑的手腕运动,SPARC分数分别为-5.18,而基线为-5.56和-6.05。

英文摘要

Human egocentric video is a scalable supervision source for humanoid policy learning, but current pipelines struggle with hand-object occlusion, oversimplified motion, or specialized capture hardware. We introduce AgenticFocus, a Mixed Reality synthesis pipeline that converts ordinary first-person-view human videos into robot-trainable demonstrations by restoring occluded object geometry, reconstructing full-hand motion, and retargeting it to a humanoid embodiment through camera-relative alignment and layered compositing. The resulting dataset pairs focused visual observations with synchronized robot actions and states. AgenticFocus achieves lower trajectory error and smoother wrist motion than cross-embodiment baselines, with SPARC scores of -5.18 versus -5.56 and -6.05.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑