强化学习中基于不确定性驱动的回放记忆
Uncertainty-Driven Replay Memory for Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本研究提出一种名为不确定性驱动的回放记忆(UDRM)的经验回放缓冲区,通过基于不确定性估计更新存储的记忆,使RL智能体在训练中获得更高奖励、提升泛化能力。
中文摘要 AI 辅助
不确定性估计为强化学习(RL)智能体提供了重要能力,值得注意的是,估计不确定性可减少训练时间,并使智能体通过利用与动作是否有助于探索环境中已知部分与相对未知部分相关的信息,随时间获得更高的奖励。在这项工作中,我们提出了一种RL中常用的经验回放缓冲区的新形式,称为不确定性驱动的回放记忆(UDRM),该方法基于RL模型在训练过程中获得的不确定性估计,对内部存储的记忆采用更新方案。与现有RL形式(通常使用时间差误差或转移分布来更新回放记忆缓冲区并训练RL控制器)不同,我们的方案使缓冲区偏向存储更多不确定的转移,这将在整个训练过程中提高RL智能体的泛化能力。实验结果表明,与其他现有的不确定性感知RL框架相比,我们提出的不确定性感知回放缓冲区使RL智能体在训练期间能获得更高的奖励。
英文摘要
Uncertainty estimation provides promising capabilities for reinforcement learning (RL) agents. Notably, estimating uncertainty can reduce the training time and enable agents to obtain greater rewards over time by exploiting information related to whether an action would facilitate exploration of portions of an environment that are well-known versus those that are relatively unknown. In this work, we propose a novel formulation of the experience replay buffer commonly used in RL that we call uncertainty-driven replay memory (UDRM), which entails an update scheme for internally stored memories based on uncertainty estimates obtained by an RL model during training. In contrast to existing forms of RL, which typically use temporal difference error or the distribution of transitions to update the replay memory buffer and train RL controllers, our scheme biases the memory buffer to store more uncertain transitions that will improve an RL agent's generalization throughout training. Experimental results demonstrate that our proposed uncertainty-aware replay buffer enables an RL agent to obtain higher rewards during training compared to other existing uncertainty-aware RL frameworks.
发表机构
- Rochester Institute of Technology(罗切斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。