发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人在回路强化学习利用人类经验不足和模仿惩罚限制的问题,提出ReF-HIL框架,通过人类参考引导的价值塑造和人类动作围栏,在5个真实操作任务上18-63分钟达到90%自主成功率,最终成功率91.7-100%。
AI 中文摘要
人在回路强化学习(HIL-RL)通过将自主学习与人类演示和在线修正相结合,为高效训练机器人操作策略提供了一条有前景的途径。然而,在价值学习中未能充分利用成功的人类经验会延长昂贵的现实世界训练时间,而持续的模仿惩罚可能限制价值驱动的策略改进。为解决这些局限性,我们提出了ReF-HIL,一种高效的人机交互强化学习框架,利用人类指导加速学习过程。人类参考引导的价值塑造从成功的人类经验中学习一个独立的价值参考,以指导在线价值学习,同时纳入局部修正反馈。人类动作围栏定义了一个学习到的人类动作邻域,允许在邻域内进行价值驱动的优化以获得更好的性能而无需模仿惩罚,同时在邻域外约束策略和价值更新。在五个多样且具有挑战性的现实世界操作任务上的实验表明,与评估的基线相比,整体学习效率和成功率均有所提高。具体而言,ReF-HIL在仅18-63分钟的主动训练内达到90%的自主成功率,并实现91.7-100%的最终成功率。这些结果突显了人类引导的强化学习在现实世界中高效获取可靠操作技能的潜力。项目网站:此https URL。
英文摘要
Human-in-the-loop reinforcement learning (HIL-RL) offers a promising route to efficient training of robotic manipulation policies by combining autonomous learning with human demonstrations and online corrections. However, insufficient use of successful human experience in value learning prolongs costly real-world training, while persistent imitation penalties can limit value-driven policy improvement. To address these limitations, we propose ReF-HIL, an efficient HIL-RL framework that uses human guidance to accelerate the learning process. Human-Reference-Guided Value Shaping learns an independent value reference from successful human experience to guide online value learning, while incorporating local corrective feedback. A Human Action Fence defines a learned human-action neighborhood, allowing value-driven optimization for better performance without imitation penalties inside while constraining policy and value updates outside. Experiments on five diverse and challenging real-world manipulation tasks demonstrate improved overall learning efficiency and higher success rates compared with the evaluated baselines. Specifically, ReF-HIL reaches 90% autonomous success in only 18-63 minutes of active training and achieves final success rates of 91.7-100%. These results highlight the potential of human-guided reinforcement learning to acquire reliable manipulation skills efficiently in the real world. Project website: https://anonymous.4open.science/w/ReF-HIL-7762/
Comments8 pages, 7 figures. Shaoyin Luo and Song Wang contributed equally to this work