arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于灵巧操作的极简重定向引导强化学习方法

A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

Yunhai Feng, Natalie Leung, Jiaxuan Wang, Lujie Yang, Haozhi Qi, Preston Culbertson

arXiv 2607.11874首次发表:更新:

发表机构

Cornell University; Amazon FAR(康奈尔大学; 亚马逊FAR)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何将重定向引导强化学习用于灵巧操作,提出REGRIND方法,从单人演示学习策略,经重定向、模拟训练、零样本转移到硬件,在丰富接触任务中产生类人行为,还分析了模拟到现实转移的关键因素。

AI 中文摘要

近期人形机器人全身控制的工作通过简单方法取得成功:将人类运动重定向到机器人运动学参考,然后通过强化学习训练策略来跟踪。但该方法在灵巧操作中如何应用并不明确,因操作涉及复杂动力学。我们提出REGRIND,一种极简重定向引导强化学习流程,从单人演示中学习灵巧操作策略。它将人类手部与物体运动重定向到保留手部与物体空间及接触关系的机器人参考,在模拟中训练残差强化学习策略以跟踪沿该参考的以物体为中心的关键点,并通过仔细的系统识别将策略零样本转移到硬件。在丰富接触的工具使用任务中,该策略在两种不同多指手上产生流畅、类人行为。通过系统硬件实验,识别并分析了灵巧操作中模拟到现实转移的关键因素,为丰富接触场景下基于重定向的学习提供实用指导。

英文摘要

Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.

CommentsWebsite: https://yunhaifeng.com/REGRIND

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑