arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于接触锚定重定向和残差策略学习的人类示教灵巧机器人操作

Dexterous Robot Manipulation from Human Demonstrations via Contact-Anchored Retargeting and Residual Policy Learning

Zihao Yang, Chengyuan Liu, Yu Zhou, Runze Lv, Tianyu Cui, Sheng Yi, Haohua Zhu, Irvine Lu, JieQ Sun

arXiv 2609.24093首次发表:更新:

发表机构

DexGEM Lab; Shanghai Jiao Tong University; Tongji University; DexRobot Co. Ltd.(DexGEM实验室; 上海交通大学; 同济大学; DexRobot有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对人类示教缺乏接触力导致灵巧操作学习困难的问题,提出物理细化、接触锚定重定向和残差策略学习三阶段流水线,无需真实机器人数据即可将人类运动捕捉转化为机器人策略,显著提升成功率并跨手迁移。

AI 中文摘要

从示教中学习灵巧操作受到数据的瓶颈限制:决定抓取是否成功的接触力在每一个可扩展的人类示教来源中都是缺失的。本文基于两个观察。第一,从人手到机器人手转换后保留下来的是示教的接触结构——哪些手指区域接触哪些物体位置,以及接触顺序——而非关节运动。第二,物理一致性无需针对每个任务进行工程化设计:一个单一的残差强化学习(RL)策略,在多样化的示教上训练一次,即可将运动学记录修复为物理一致、带接触标注的轨迹,并且相同的残差公式在重定向后恢复动态可行性。这些观察产生了一个三阶段流水线,将人类运动捕捉记录转换为灵巧机器人策略,且无需真实机器人训练数据:使用模拟MANO手进行物理细化以恢复接触和力,接触锚定重定向通过独立于手形态的目标函数转移示教的接触结构,残差策略学习将结果适应于机器人驱动。该流水线重建了25,454条单手轨迹(成功率从7.3%提升至59.3%)和25个双手任务(从16.0%提升至62.4%),每个设置使用一个共享策略,将一个人类数据集迁移到四个形态不同的机器人手(提升62.4个百分点),并在物理硬件上执行四个接触丰富的双手任务,且零真实机器人训练数据。

英文摘要

Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine whether a grasp succeeds are absent from every scalable source of human demonstrations. This paper builds on two observations. First, what survives the change from a human hand to a robot hand is the contact structure of a demonstration - which finger regions touch which object locations, and in what order - rather than its joint motion. Second, physical consistency need not be engineered per task: a single residual reinforcement learning (RL) policy, trained once across diverse demonstrations, can repair kinematic recordings into physically consistent, contact-annotated trajectories, and the same residual formulation restores dynamic feasibility after retargeting. These observations yield a three-stage pipeline that converts human motion-capture recordings into dexterous robot policies with no real-robot training data: physics refinement with a simulated MANO hand recovers contacts and forces, contact-anchored retargeting transfers the demonstrated contact structure through an objective independent of hand morphology, and residual policy learning adapts the result to robot actuation. The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.

Comments16 pages, 6 figures, 5 tables. Technical report. Code: https://github.com/DexGEM-Lab/real2sim2real

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑