将新颖性与意外性结合用于基于图像的强化学习中的经验优先级排序与探索
Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
该研究提出NSPER及扩展的NSPER+R,将新颖性与意外性结合,在DeepMind Control Suite任务上提升了基于图像的强化学习的训练效率与收敛速度。
中文摘要 AI 辅助
样本效率是强化学习(RL)的核心挑战,尤其在基于图像的领域中,智能体必须从高维视觉输入中学习。传统采样常依赖随机或次优的经验选择,导致冗余更新与学习缓慢。提升效率需要同时具备对信息丰富经验进行优先级排序并鼓励有效探索的机制。优先经验回放(PER)通过重用高价值转换解决了部分挑战,而内在奖励则促进对新颖或不确定状态的探索,但二者的整合尚未得到广泛研究。本文提出新颖性与意外性优先经验回放(NSPER),其利用新颖性捕捉代表性不足的状态,利用意外性暴露智能体对环境理解的缺口。我们进一步扩展出NSPER+R,将这些信号作为内在奖励整合,共同提升回放质量与探索效果。在DeepMind Control Suite任务上的实验表明,NSPER与NSPER+R相较于现有基于图像的RL方法,提升了训练效率与收敛速度。
英文摘要
Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from high-dimensional visual inputs. Traditional sampling often relies on random or suboptimal experience selection, leading to redundant updates and slow learning. Improving efficiency requires mechanisms that prioritize informative experiences while also encouraging effective exploration. Prioritized Experience Replay (PER) addresses part of this challenge by reusing high-value transitions, while intrinsic rewards promote the exploration of novel or uncertain states. However, their integration has not been extensively studied. This paper introduces Novelty and Surprise Prioritized Experience Replay (NSPER), which uses novelty to capture underrepresented states and surprise to expose gaps in the agent's understanding of the environment. We further extend this with NSPER+R, integrating these signals as intrinsic rewards to jointly improve replay quality and exploration. Experiments on DeepMind Control Suite tasks show that NSPER and NSPER+R improve training efficiency and convergence speed compared to existing methods in image-based RL.
发表机构
- University of Auckland(奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。