发表机构
Korea Advanced Institute of Science and Technology(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对强化学习中稀疏与延迟奖励的探索难题,本文提出ENTINEX方法,利用熵信息识别状态分布边界并分配内在奖励,实验表明其性能优于现有方法。
AI 中文摘要
在强化学习中,由于引导学习过程的反馈有限,具有稀疏和延迟奖励的探索是一个重大挑战。解决该问题需要在状态空间中进行广泛探索以发现有价值的奖励信号。本文提出一种名为探索用熵信息(ENTINEX)的新方法,该方法通过激励智能体探索状态分布边界之外来增强探索,ENTINEX通过为这些边界分配内在奖励、利用熵信息有效识别边界来实现这一点。通过大量实验,我们证明ENTINEX在具有稀疏和延迟奖励特征的环境中持续提升探索性能,实验结果显示ENTINEX优于现有探索方法,凸显其在稀疏和延迟奖励场景中的有效性。
英文摘要
In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process. Addressing this issue requires extensive exploration in the state space to discover valuable reward signals. In this paper, we propose Entropic Information for Exploration (ENTINEX), a novel method that enhances exploration by incentivizing agents to explore beyond the boundaries of the state distribution. ENTINEX achieves this by assigning intrinsic rewards to these boundaries, leveraging entropic information to identify them effectively. Through extensive experimentation, we demonstrate that ENTINEX consistently improves exploration performance in environments characterized by sparse and delayed rewards. Our experimental results show that ENTINEX outperforms existing exploration methods, highlighting its effectiveness in both sparse and delayed reward scenarios.