arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用人类演示的接触力引导学习灵巧操作

Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration

Xinghao Zhu, Zixi Liu, Shalin Jain, Chenran Li, Milad Noori, Michael Andres Lin, Huihua Zhao, John Welsh, Mrinal Verghese, Wei Liu, Tingwu Wang, Xingye Da, Zhengyi Luo, Vishal Kulkarni, Naema Bhatti, Yuke Zhu, Linxi Fan, Bowen Wen, Danfei Xu, Soha Pouya, Yan Chang

arXiv 2607.00033首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CHORD框架,通过物体中心的接触力空间引导,利用人类演示的力和力矩信息增强强化学习,在包含4739个双臂灵巧操作任务的大规模基准上平均成功率82.12%,并泛化到全身操作和真实世界。

AI 中文摘要

灵巧机器人操作可以从丰富的人类演示中受益,但将这些演示转化为机器人策略仍然具有挑战性。我们提出了CHORD(来自人类演示的机器人灵巧操作中的接触力引导),这是一个用于刚体和铰接物体长时域操作的强化学习框架。关键思想是以物体为中心的接触力空间引导:我们通过人类和机器人运动在物体上产生的力和力矩来表示这些运动,从而可以通过诱导的瞬时运动来度量相似性。这种引导使得强化学习在接触丰富的灵巧操作中更具可扩展性。我们进一步引入了一个大规模仿真基准,包含从动作捕捉数据集和内部重建视频构建的4739个双臂灵巧操作任务。在1831个基准任务上的评估中,CHORD达到了82.12%的平均成功率,展示了强大的可扩展性。CHORD还从仅手部和第三人称演示泛化到全身操作,成功率达到90.77%,并且学习到的策略在开环和闭环设置下都能迁移到真实世界。

英文摘要

Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑