层次行为空间
Hierarchical Behaviour Spaces
- Meta Superintelligence Labs(Meta超智能实验室)
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出层次行为空间(HBS)方法,通过线性组合奖励函数诱导行为空间,提升策略表达能力,在NetHack环境中验证了其性能,发现层级优势源于探索而非长期推理。
AI中文摘要:
最近的层次强化学习研究显示,当学习一组预定义的选项奖励函数时,可以扩展到数十亿时间步。我们展示,替代使用每个选项一个奖励函数,奖励函数可以有效诱导行为空间,通过让控制器指定奖励函数的线性组合,允许更丰富的策略集被表示。我们称之为层次行为空间(HBS)。我们在NetHack学习环境中评估了HBS,展示了强劲的性能。我们进行了一系列实验,并确定,或许与传统智慧相反,我们方法中的层级优势源于增加的探索而非长期推理。
英文摘要:
Recent work in hierarchical reinforcement learning has shown success in scaling to billions of timesteps when learning over a set of predefined option reward functions. We show that, instead of using a single reward function per option, the reward functions can be effectively used to induce a space of behaviours, by letting the controller specify linear combinations over reward functions, allowing a more expressive set of policies to be represented. We call this method Hierarchical Behaviour Spaces (HBS). We evaluate HBS on the NetHack Learning Environment, demonstrating strong performance. We conduct a series of experiments and determine that, perhaps going against conventional wisdom, the benefits of hierarchy in our method come from increased exploration rather than long term reasoning.