目标条件双动作模仿学习用于灵巧双臂机器人操作
Goal-conditioned dual-action imitation learning for dexterous dual-arm robot manipulation
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对长时程灵巧操作可变形物体(如剥香蕉)的难题,提出目标条件双动作模仿学习方法,通过局部与全局动作结合防止复合误差,并在真实双臂机器人上成功完成剥香蕉任务。
AI中文摘要:
长时程灵巧机器人操作可变形物体(如剥香蕉皮)是一项困难的任务,原因在于物体建模的困难以及缺乏关于稳定和灵巧操作技能的知识。本文提出了一种目标条件双动作(GC-DA)深度模仿学习(DIL)方法,该方法可以利用人类演示数据学习灵巧操作技能。以往的DIL方法将当前感官输入与反应性动作直接映射,由于模仿学习中由动作的循环计算引起的复合误差,这种方法常常失败。所提出的方法仅在需要对目标物体进行精确操作时预测反应性动作(局部动作),而在不需要精确操作时生成整个轨迹(全局动作)。这种双动作公式通过基于轨迹的全局动作有效防止了模仿学习中的复合误差,同时在反应性局部动作期间能够响应目标物体的意外变化。所提出的方法在真实双臂机器人上进行了测试,并成功完成了剥香蕉皮任务。本文及相关工作的数据可在以下网址获取:https://sites.google.com/view/multi-task-fine。
英文摘要:
Long-horizon dexterous robot manipulation of deformable objects, such as banana peeling, is a problematic task because of the difficulties in object modeling and a lack of knowledge about stable and dexterous manipulation skills. This paper presents a goal-conditioned dual-action (GC-DA) deep imitation learning (DIL) approach that can learn dexterous manipulation skills using human demonstration data. Previous DIL methods map the current sensory input and reactive action, which often fails because of compounding errors in imitation learning caused by the recurrent computation of actions. The method predicts reactive action only when the precise manipulation of the target object is required (local action) and generates the entire trajectory when precise manipulation is not required (global action). This dual-action formulation effectively prevents compounding error in the imitation learning using the trajectory-based global action while responding to unexpected changes in the target object during the reactive local action. The proposed method was tested in a real dual-arm robot and successfully accomplished the banana-peeling task. Data from this and related works are available at: https://sites.google.com/view/multi-task-fine.