arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36473cs.AIcs.LG

超越状态到达:面向内在动机选项发现的学习抽象

Going Beyond State-Reaching: Learning Abstractions for Intrinsically Motivated Option Discovery

  • Amazon(亚马逊)
  • Brown University(布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

Akhil Bagaria, Anita De Mello Koch, George Konidaris

AI总结:

提出一种为每个子目标识别相关特征子集的选项发现算法,学习抽象可迁移选项,在稀疏奖励图像领域加速探索。

AI中文摘要:

通过选项进行时间抽象可以改善大型环境中的探索。然而,现有的选项发现算法寻找的子目标同时针对状态的所有方面。这种状态到达方法产生的选项仅适用于状态空间的狭窄区域,最终导致选项数量爆炸,使智能体不堪重负,并阻碍其首要任务——奖励最大化——的进展。我们提出一种算法,该算法为每个子目标识别一小部分相关特征子集,从而产生广泛泛化并加速探索的选项。我们的方法学习抽象、可迁移的选项,并在三个稀疏奖励、基于图像的领域(包括Atari游戏MontezumasRevenge)中实现快速探索。

英文摘要:

Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.

↑