FlashDexRetarget:通过多动作重定向加速灵巧操作数据生成
FlashDexRetarget: Accelerating Dexterous Manipulation Data Generation through Multi-Motion Retargeting
- KAIST AI(韩国科学技术院人工智能学院)
- Holiday Robotics
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
FlashDexRetarget提出基于强化学习的灵巧动作重定向框架,结合点云观测与互补奖励,在50动作基准上达90%成功率,训练计算量较基线降低百倍。
AI中文摘要:
人类手-物体演示为灵巧机器人操作数据提供了可复用的来源,但跨实体迁移这些数据需要物理可行的重定向。现有的基于物理的方法在重定向成功率、动作特定训练效率或两者兼有方面面临局限。为解决这些限制,我们提出了FlashDexRetarget,一个基于强化学习的框架,用于高成功率、高效的灵巧动作重定向。为使演示的交互更易学习,我们将物体点云观测、手-物体距离特征和未来轨迹编码与互补奖励相结合,这些奖励监督物体运动并参考手-物体关系。为进一步加速学习,我们采用独立的左手和右手演员-评论家网络,并将离策略算法FlashSAC适配到灵巧动作跟踪。在涵盖单物体和双物体交互的50个动作基准上,FlashDexRetarget实现了90%的成功率,约为所评估的基于采样的基线成功率的2.5倍,同时所需的训练计算量比所评估的基于强化学习的基线少高达100倍。在XHand和Sharpa Wave Hand上的评估均显示出一致的改进,组件消融研究检验了我们设计选择的贡献。在50个动作基准之外,使用200、500和1000个动作的实验表明,我们的方法在更大规模下保持稳定,并随着训练集的增长更高效地产生成功的重定向动作。使用真实世界捕获演示的定性回放结果进一步说明了我们的框架对记录的人类操作的适用性。视频和代码可在以下网址获取:此https URL
英文摘要:
Human hand-object demonstrations provide a scalable source of data for dexterous robot learning, but transferring them across embodiments requires physically feasible retargeting. Existing physics-based methods typically optimize each demonstration independently, leading to either limited success under finite simulation budgets or training costs that grow with dataset size. We introduce FlashDexRetarget, an RL framework for multi-reference dexterous retargeting. We formulate retargeting as multi-reference tracking, jointly learning a single policy across many demonstrations with off-policy RL and geometric supervision of the demonstrated interactions. This shared training formulation amortizes optimization across references while enabling the policy to track diverse hand-object interactions. On a 50-motion benchmark from TACO, OakInk2, and HOT3D using XHand and Sharpa Wave Hand as target embodiments, FlashDexRetarget retargets 90% of demonstrations using about 30 GPU-hours, compared with about 46% at about 3,000 GPU-hours for CHORD. This corresponds to about 100 times lower training compute and a 44-percentage-point improvement in retargeting success. Ablations examine the key design choices, while experiments with up to 1,000 motions and real-world replay further demonstrate the scalability and practical applicability of our method.