发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究无模型强化学习中规划作为涌现行为的现象,发现神经架构的隐藏状态结构是决定因素,关系隐藏状态网络能获规划机制,恢复环境转移结构并改进策略,解释了涌现规划现象并引发相关问题。
AI 中文摘要
强化学习通常分为基于模型和无模型方法。在这种分类中,基于模型的方法在学习到的世界模型上进行前瞻性规划,而无模型方法学习反应式状态-动作映射。然而,最近的工作表明规划可以仅从无模型强化学习中涌现。到目前为止,这种行为从纯粹的奖励最大化目标中出现的条件仍不清楚。本文提供证据表明,在观察到的案例中,神经架构的隐藏状态结构是决定因素。我们发现一个关系隐藏状态网络,每个状态锚定到一个环境状态并沿学习到的关系交换消息,获得了一种规划机制。这些隐藏状态在其学习到的关系中恢复环境的转移结构,并在决策时通过在学习到的图上进行规划来改进策略。在一个必须额外发现哪些单元格代表哪些状态的匹配控制代理中,不会出现这种绑定,也不会由此产生规划。我们认为这解释了在无模型强化学习中观察到的涌现规划现象,并提出了这种涌现规划在更普遍情况下可能有多常见的问题。最后,我们假设所发现的机制可以描述规划如何通过神经架构先验从人类大脑中的纯粹奖励最大化中涌现。
英文摘要
Reinforcement learning is conventionally divided into model-based and model-free methods. In this taxonomy, model-based methods perform lookahead planning over a learned world model, whereas model-free methods learn a reactive state-action mapping. Recent work, however, has shown that planning can emerge from model-free reinforcement learning alone. The conditions under which this behavior emerges from a pure reward-maximization objective have so far remained unclear. In this paper, we present evidence that, in the observed cases, the hidden-state structure of the neural architecture is the deciding factor. We find that a network of relational hidden states, each anchored to an environment state and exchanging messages along learned relations, acquires a planning mechanism. These hidden states recover the environment's transition structure in their learned relations, and improve the policy at decision time by planning over the learned graph. In a matched control agent that must additionally discover which cells represent which states, no such binding arises, and no planning follows from it. We argue that this explains the observed phenomenon of emergent planning in model-free reinforcement learning and raises the question of how common such emergent planning might be more generally. Finally, we hypothesize that the discovered mechanism could describe how planning emerges from pure reward maximization in the human brain through a neural architectural prior.