发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出自适应干预智能体UniIntervene++,通过半马尔可夫决策过程学习控制分配,动态决定干预时机与方式,在真实世界操作任务中成功率89.67%,干预率仅0.77%。
AI 中文摘要
在线强化学习(RL)使机器人策略能够通过物理交互得到改进,但随着其能力的演进,所需的辅助也会发生变化。基于离线估计或固定决策规则的现有干预策略因此可能与当前策略不匹配。为解决这一问题,我们提出了UniIntervene++,一种自适应干预智能体,它在在线RL过程中学习在自主执行和异构辅助行为之间分配控制权。具体而言,UniIntervene++首先将演进的RL策略、轨迹修正以及任务结构化的CodePolicy作为选项(Options)统一到一个半马尔可夫决策过程中,并在线学习它们的相对价值。在此基础上,能力自适应干预通过无辅助执行定期探测RL策略,使控制分配对其演进的能力保持响应。最后,耦合经验学习使得辅助行为能够改进RL策略,而策略演进的成果反过来又会重塑未来的干预决策。通过这种方式,UniIntervene++在RL策略改进过程中联合决定何时干预、如何干预以及何时归还控制权。在五个真实世界的操作任务中,UniIntervene++实现了89.67%的平均成功率,比所有基线高出至少6个百分点,同时将人工干预降低至0.77%,相对于最佳基线至少相对减少94.6%。代码可在我们的GitHub仓库中获取。
英文摘要
Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or fixed decision rules can therefore become mismatched to the current policy. To address this, we propose UniIntervene++, an adaptive intervention agent that learns to allocate control between autonomous execution and heterogeneous assisted behaviors during online RL. Specifically, UniIntervene++ first formulates the evolving RL policy, trajectory correction, and a task-structured CodePolicy as Options in a unified semi-Markov decision process and learns their relative values online. Building on this, competence-adaptive intervention periodically probes the RL policy through unassisted execution, keeping control allocation responsive to its evolving capability. Finally, coupled experience learning allows assisted behaviors to improve the RL policy, whose evolving outcomes in turn reshape future intervention decisions. In this way, UniIntervene++ jointly determines when to intervene, how to intervene, and when to return control as the RL policy improves. Across five real-world manipulation tasks, UniIntervene++ achieves an average success rate of 89.67%, outperforming all baselines by at least 6 percentage points, while reducing human intervention to 0.77%, a relative reduction of at least 94.6% from the best baseline. Code is available in our \href{https://github.com/dannyyudong/An-Adaptive-Intervention-Agent-for-Efficient-Real-World-Reinforcement-Learning}{GitHub repository}.
CommentsYudong Lin and Haoyuan Deng contributed equally. Ziwei Wang is the corresponding author. Code is available in our \href{https://github.com/dannyyudong/An-Adaptive-Intervention-Agent-for-Efficient-Real-World-Reinforcement-Learning}{GitHub repository}