发表机构
Nara Institute of Science and Technology (NAIST); University of Trento(奈良先端科学技术大学院大学; 特伦托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人机具身差异导致演示不可行的问题,提出基于经验的可行性感知生成对抗模仿学习(EF-GAIfO),动态扩展可行区域,并在模拟运动任务和真实四足机器人抓取任务中验证有效性。
AI 中文摘要
随着越来越多无需机器人的演示界面被使用,这些界面提供无动作标签的状态轨迹,从观察中模仿已成为从人类演示中学习机器人行为的一种有前景的方法。然而,由于人类与机器人之间在具身和动力学上的差异,演示的人类动作可能对机器人不可行,从而可能降低策略性能。在本研究中,我们提出了基于经验的可行性感知生成对抗模仿学习(EF-GAIfO),该方法从机器人自身的经验中估计仅状态演示的可行性,而不是依赖显式动力学模型或大型先验探索数据集。EF-GAIfO的一个关键特征是,可行性的概念随策略学习而演变:随着策略改进和机器人经历更广泛的状态转换,可行区域逐渐扩大,允许将更多演示纳入学习。这使得可行性感知模仿能够适应当前策略学习阶段,而不是依赖预先设计的可行性标准。我们在模拟中的运动任务和真实四足机器人执行物体到达与抓取任务上验证了EF-GAIfO的有效性。
英文摘要
With the increasing use of robot-free demonstration interfaces that provide state trajectories without action labels, imitation from observation has become a promising approach for learning robot behaviors from human demonstrations. However, due to differences in embodiment and dynamics between humans and robots, demonstrated human motions may not be feasible for the robot, potentially degrading policy performance. In this study, we propose Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO), which estimates the feasibility of state-only demonstrations from the robot's own experience rather than relying on explicit dynamics models or large prior exploration datasets. A key feature of EF-GAIfO is that the notion of feasibility evolves with policy learning: as the policy improves and the robot experiences a broader range of state transitions, the feasible region is progressively expanded, allowing additional demonstrations to be incorporated into learning. This enables feasibility-aware imitation that adapts to the current stage of policy learning, rather than relying on a pre-designed feasibility criterion. We validate the effectiveness of EF-GAIfO on a locomotion task in simulation and on a real quadruped robot performing a object-reaching-and-grasping task.