arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

内在奖励何时导致探索?

When Do Intrinsic Rewards Lead to Exploration?

Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

arXiv 2610.02159首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出反事实信息作为探索标准,证明现有内在奖励在简单环境中获取反事实信息时是帕累托次优的,并给出成功条件及改进目标。

AI 中文摘要

内在奖励旨在通过为智能体的经验赋予价值来引导强化学习中的探索,例如通过预测误差或学习进度。然而,最大化这些奖励不一定能产生最有信息量的经验。我们提出了一种正式的探索标准,通过比较策略所获取的反事实信息来评估策略:即它们的历史在替代策略下能够替代经验的程度。我们构建了一个简单环境,在该环境中,指定的基于计数、预测误差、赋能和信息增益的目标,其最大化策略在获取反事实信息方面是帕累托次优的。我们解释了这些失败,并建立了现有内在奖励成功促进最优探索的条件。我们还构建了一个目标,该目标在探索根据我们的标准严格改进时赋予更高的价值。

英文摘要

Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.

Comments45 pages, 4 figures; includes mathematical appendices. Code, data, and Lean proof sources: https://github.com/scottviteri/what-is-exploration

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑