超越状态到达:面向内在动机选项发现的学习抽象
Going Beyond State-Reaching: Learning Abstractions for Intrinsically Motivated Option Discovery
- Amazon(亚马逊)
- Brown University(布朗大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出一种为每个子目标识别相关特征子集的选项发现算法,学习抽象可迁移选项,在稀疏奖励图像领域加速探索。
AI中文摘要:
通过选项进行时间抽象可以改善大型环境中的探索。然而,现有的选项发现算法寻找的子目标同时针对状态的所有方面。这种状态到达方法产生的选项仅适用于状态空间的狭窄区域,最终导致选项数量爆炸,使智能体不堪重负,并阻碍其首要任务——奖励最大化——的进展。我们提出一种算法,该算法为每个子目标识别一小部分相关特征子集,从而产生广泛泛化并加速探索的选项。我们的方法学习抽象、可迁移的选项,并在三个稀疏奖励、基于图像的领域(包括Atari游戏MontezumasRevenge)中实现快速探索。
英文摘要:
Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.