Res-HIL:面向样本高效灵巧操作的人类引导残差强化学习
Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
- Siemens AG, Research and Predevelopment(西门子股份公司,研究与预开发部)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
Res-HIL提出人在回路残差强化学习框架,在冻结的模仿策略上学习纠正动作,结合直接监督与奖励塑形,仅用20个示范和十分钟训练即在五个灵巧操作任务上超越现有方法。
中文摘要 AI 辅助
模仿学习使机器人能够从示范中获取操作技能,但由此产生的策略在训练数据之外可能失效,而收集更多示范需要大量人力投入。人在回路强化学习在在线训练期间利用纠正性反馈,但通常学习完整的任务策略,而非优化预训练的模仿策略。我们提出Res-HIL,一种人在回路残差强化学习框架,在冻结的模仿策略之上学习纠正动作。每次人类干预提供两种互补的学习信号:对残差策略的直接监督以及对先前自主行为的奖励塑形。Res-HIL将这些信号与残差策略的零初始化相结合,以稳定并加速在线学习。我们在五个接触丰富的操作任务上评估Res-HIL,这些任务涵盖高精度和长时程行为。仅使用20个初始示范,经过十分钟在线训练后,Res-HIL在每项任务上均优于最先进的全策略人在回路强化学习和无人类指导的残差微调方法。Res-HIL改进了其预训练基础策略,并优于使用五倍示范训练的模仿策略。消融研究表明,直接残差监督对性能至关重要,而干预感知的奖励塑形显著提高了训练效率。
英文摘要
Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.