arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15707cs.RO

GAINS:在强化学习中利用不一致的人类干预信号

GAINS: Leveraging Inconsistent Human Intervention Signals in Reinforcement Learning

Xinyi Zhang, Yinuo Zhao, Pei Ren, Lechun Jiang, Huiqian Jin, Lei Sun, Dapeng Wu, Zhengping Che, Chi Harold Liu, Jian Tang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出GAINS框架,采用带分位数Q网络的分布强化学习建模人类干预引发的回报变异性,结合悲观探索策略,在模拟与现实场景中使任务成功率较RLIF高22%,故障恢复成功率最高提升43%,凸显该建模对现实部署的重要性。

中文摘要 AI 辅助

通过人类干预修正机器人操控策略对现实世界部署具有重要意义,但人类操作者在提供的动作和干预信号时机上天生存在缺陷。前者已在强化学习(RL)中被广泛研究,而后者却未得到充分探索。在高控制频率下,人类干预信号常存在延迟且随时间和状态空间呈现不一致性。本研究提出GAINS,一种用于在强化学习中利用不一致人类干预信号的框架。GAINS的核心是采用带分位数Q网络的分布强化学习,以建模由稀疏任务奖励和不一致人类干预引发的回报变异性。基于该分布表示,我们引入一种悲观探索策略,以在人类修正下促进安全且样本高效的学习。我们在四个不同的模拟操控任务和两个具有挑战性的现实场景中,与最先进的基于干预的方法对GAINS进行评估。GAINS的任务成功率比RLIF高22%,在故障场景中的恢复成功率提升高达43%。这些结果凸显了对人类缺陷引发的回报变异性进行建模,对基于干预的学习在现实世界部署的重要性。

英文摘要

Correcting robot manipulation policies through human intervention holds great promise for real-world deployment, yet human operators are inherently imperfect in both the actions they provide and the timing of their intervention signals. While the former has been extensively discussed in reinforcement learning (RL), the latter remains underexplored. At high control frequencies, human intervention signals are often delayed and inconsistent across time and state space. In this work, we present GAINS, a framework for leveraging inconsistent human intervention signals in RL. At the core of GAINS, we employ distributional RL with quantile Q-networks to model the return variability induced by sparse task rewards and inconsistent human interventions. Building on this distributional representation, we introduce a pessimistic exploration strategy that promotes safe and sample-efficient learning under human corrections. We evaluate GAINS on four diverse simulated manipulation tasks and two challenging real-world scenarios against state-of-the-art intervention-based methods. GAINS achieves a 22% higher task success rate than RLIF and improves recovery success by up to 43% in failure scenarios. These results highlight the importance of modeling return variability induced by human imperfection for real-world deployment of intervention-based learning.

发表机构

  • Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)

机构由 AI 辅助整理,请以论文原文为准。

↑