arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.09013cs.ROcs.CV

通过3D手-物体轨迹重建从RGB人类视频中学习灵巧操作策略

Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction

  • Carnegie Mellon University(卡内基梅隆大学)
  • Georgia Institute of Technology(佐治亚理工学院)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Hongyi Chen, Tony Dong, Tiancheng Wu, Liquan Wang, Yash Jangir, Yaru Niu, Yufei Ye, Homanga Bharadhwaj, Zackory Erickson, Jeffrey Ichnowski

更新

AI总结:

VIDEOMANIP通过3D手-物体轨迹重建从RGB视频直接学习灵巧操作策略,实现无需设备的高效训练和高成功率抓取

AI中文摘要:

多指机械手操作和抓取因高维动作空间和获取大规模训练数据困难而具有挑战性。现有方法大多依赖于佩戴设备或专用传感设备的人类远程操作来捕捉手-物体交互,这限制了可扩展性。在本工作中,我们提出VIDEOMANIP,一种无需设备的框架,直接从RGB人类视频中学习灵巧操作。利用计算机视觉的最新进展,VIDEOMANIP通过估计人体手部姿态、物体网格并重新定向重建的人类动作到机械手,从单目视频中重建显式的3D机器人-物体轨迹。为了使重建的机器人数据适合灵巧操作训练,我们引入了手-物体接触优化与以交互为中心的抓取建模,以及一种演示合成策略,该策略从单个视频生成多样化的训练轨迹,使政策学习具有泛化能力而无需额外的机器人演示。在模拟中,学习的抓取模型在20种不同物体上实现了70.25%的成功率。在现实世界中,从RGB视频训练的操作策略在七个任务上平均实现了62.86%的成功率,优于基于重定向的方法,高出15.87%。项目视频可在videomanip.github.io上获取。

英文摘要:

Multi-finger robotic hand manipulation and grasping are challenging due to the high-dimensional action space and the difficulty of acquiring large-scale training data. Existing approaches largely rely on human teleoperation with wearable devices or specialized sensing equipment to capture hand-object interactions, which limits scalability. In this work, we propose VIDEOMANIP, a device-free framework that learns dexterous manipulation directly from RGB human videos. Leveraging recent advances in computer vision, VIDEOMANIP reconstructs explicit 3D robot-object trajectories from monocular videos by estimating human hand poses, object meshes, and retargets the reconstructed human motions to robotic hands for manipulation learning. To make the reconstructed robot data suitable for dexterous manipulation training, we introduce hand-object contact optimization with interaction-centric grasp modeling, as well as a demonstration synthesis strategy that generates diverse training trajectories from a single video, enabling generalizable policy learning without additional robot demonstrations. In simulation, the learned grasping model achieves a 70.25% success rate across 20 diverse objects using the Inspire Hand. In the real world, manipulation policies trained from RGB videos achieve an average 62.86% success rate across seven tasks using the LEAP Hand, outperforming retargeting-based methods by 15.87%. Project videos are available at videomanip.github.io.

↑