arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于内在好奇心的强化学习内生探索

Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

Armando Vieira

arXiv 2609.05650首次发表:更新:

发表机构

University of Tartu(塔尔图大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种由内在好奇心驱动的强化学习框架,在液态状态机上实现,通过中等不一致性促进探索,在LunarLanderv2和BipedalWalkerv3基准上取得与PPO和ICM相当的竞争力。

AI 中文摘要

我们提出了一种由内在好奇心驱动的强化学习框架,专为非平稳环境以及奖励稀疏、延迟、无信息或缺失的场景而设计。在我们的模型中,动作选择由外部奖励和一种认知动机机制共同引导,该机制使智能体偏向于结构化的探索方向。核心假设是,有效的探索出现在中等程度的不一致(incoherence)水平下,而性能在过于僵化和过于无序的动态中都会下降。为验证这一想法,我们在液态状态机(Liquid State Machine, LSM)基底上实现了该框架,并在两个标准基准上进行了评估:离散动作的LunarLanderv2和连续控制的BipedalWalkerv3。与包括近端策略优化(Proximal Policy Optimization, PPO)和内在好奇心模块(Intrinsic Curiosity Module, ICM)在内的成熟深度强化学习算法相比,所提方法在两个任务上均取得了具有竞争力的性能。我们进一步表明,在相同的分析下,Active Inference智能体并未恢复出好奇心窗口,这表明所提出的动态机制捕获了一种独特的探索模式。

英文摘要

We propose a reinforcement learning framework in which exploration is driven by intrinsic curiosity, designed for scenarios where environments are non-stationary and rewards are sparse, delayed, uninformative, or absent. In our model, action selection is guided by a combination of external rewards and an epistemic motivation mechanism that biases the agent toward structured exploratory directions. The central hypothesis is that effective exploration emerges at intermediate levels of incoherence, while performance degrades under both overly rigid and overly disordered dynamics. To test this idea, we implement the framework on top of a Liquid State Machine (LSM) substrate and evaluate it on two standard benchmarks: the discrete-action LunarLanderv2 and the continuous-control BipedalWalkerv3. The proposed method achieves competitive performance on both tasks relative to established deep RL algorithms, including Proximal Policy Optimization (PPO) and Intrinsic Curiosity Module (ICM). We further show that the curiosity window is not recovered in Active Inference agents under the same analysis, suggesting that the proposed dynamics capture a distinct exploration regime

Comments25 pages, 5 figurees

DOI:10.20944/preprints202608.1449.v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑