EgoSpeedUp:将人类操作节奏迁移至机器人策略
EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies
- National Institute of Advanced Industrial Science and Technology (AIST)(日本国立产业技术综合研究所(AIST))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EgoSpeedUp通过人类演示的节奏监督,对齐并迁移分阶段操作节奏至机器人模仿学习,在真实任务中提升成功率25个百分点并减少36.5%执行时间。
AI中文摘要:
通过模仿学习训练的机器人操作策略不仅继承了演示行为,还继承了机器人演示的保守执行节奏。现有的加速方法可以比原始演示执行得更快,但主要依据机器人侧信息或预定义的节奏因子集来确定合适的加速,这留下了如何为每个操作阶段应进展的速度获取任务适切参考的问题。我们提出EgoSpeedUp,一个利用人类操作作为机器人模仿学习时间监督的框架。我们的关键洞见是,人类演示自然揭示了任务适切、分阶段的操作节奏。给定相同任务的慢速机器人演示和人类演示,EgoSpeedUp对齐相应的操作阶段,从多个人类演示中估计它们的相对执行节奏,并通过重新定时机器人演示来迁移所得的分阶段节奏。然后,重新定时的演示用于标准行为克隆,使机器人保留其可执行的操作行为,同时学习以人类告知的节奏执行。在两个真实世界操作任务中,EgoSpeedUp将任务成功率平均提高了25个百分点(pp),同时将成功执行时间减少了36.5%。这些结果表明,人类操作节奏为学习更快、更可靠的机器人策略提供了有效的时间参考。
英文摘要:
Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning. Our key insight is that human demonstrations naturally reveal task-appropriate, phase-wise manipulation tempo. Given slow robot demonstrations and human demonstrations of the same task, EgoSpeedUp aligns corresponding manipulation phases, estimates their relative execution tempos from multiple human demonstrations, and transfers the resulting phase-wise tempo by retiming the robot demonstrations. The retimed demonstrations are then used for standard behavior cloning, allowing the robot to retain its executable manipulation behavior while learning to perform it at a human-informed tempo. Across two real-world manipulation tasks, EgoSpeedUp improves the task success rate by an average of 25 percentage points (pp) while reducing successful execution time by 36.5%. These results demonstrate that human manipulation tempo provides an effective temporal reference for learning faster and more reliable robot policies.