发表机构
Istituto Italiano di Tecnologia; University of Genova; TU Delft(意大利技术研究院; 热那亚大学; 代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对真实世界机器人操作在线强化学习的局限,提出结合CTDE与HRA的框架,在多任务实验中大幅提升了样本效率与任务成功率。
AI 中文摘要
真实世界在线强化学习(RL)为直接在物理世界中训练机器人操作策略提供了有前景的方法,避免了模拟到现实的差距,并通过人在回路交互实现策略的持续优化。近期方法通过人工干预展现出样本高效学习能力,但仍局限于较小的随机化范围,且面临多个智能体同时训练引发的非平稳性挑战。为解决这些局限,本文引入结合集中式训练与分布式执行(CTDE)及混合奖励架构(HRA)的统一框架,该框架允许多个执行器共享一个集中式多头评论者,评论者被分解为对应稀疏任务奖励的任务头和对应基于势能的抓取奖励的抓取头。据此,本文重构评论者与执行器目标,以利用分解后的Q值,同时明确考虑离散夹爪策略的类别动作分布。实验结果表明,所提框架大幅提升了样本效率与策略性能。本文在两台机械臂及一个仿真人形机器人上,针对网球与香蕉的抓取放置、锅具重置、仿真块重定位任务,在维度域随机化(其范围约为现有工作的5至25倍)下验证了所提方法。与现有最优基线相比,本文方法在网球抓取放置上的成功率从60%提升至80%,香蕉抓取放置从60%提升至90%,仿真块重定位从25%提升至95%,还成功完成了基线始终无法完成的任务。视频及更多详情可在项目网站获取:this https URL。
英文摘要
Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, avoiding the sim-to-real gap and enabling continuous policy refinement through human-in-the-loop interaction. Recent methods have demonstrated sample-efficient learning through human intervention but remain limited to small randomization ranges and encounter challenges with the non-stationarity induced by concurrently training multiple agents. To address these limitations, we introduce a unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA). This enables multiple actors to share a centralized multi-head critic. The critic is decomposed into task and grasp heads, corresponding to the sparse task reward and a potential-based grasping reward, respectively. We accordingly reformulate the critic and actor objectives to exploit the decomposed Q-values while explicitly accounting for the categorical action distribution of the discrete gripper policy. Experimental results demonstrate that the proposed framework substantially improves both sample efficiency and policy performance. We validate our approach on two robotic arms and a simulated humanoid robot across tennis ball and banana pick-and-place, pot reset, and simulated block relocation tasks under dimension-wise domain randomization, approximately 5-25x larger than those considered in prior work. Compared with a state-of-the-art baseline, our method improves the success rate from 60% to 80% on tennis ball pick-and-place, from 60% to 90% on banana pick-and-place, and from 25% to 95% on simulated block relocation, while also successfully accomplishing a task where the baseline consistently fails. Videos and more details are available at our project website: https://hil-harc.github.io/.