分层通用价值函数逼近器
Hierarchical Universal Value Function Approximators
- University of Massachusetts(马萨诸塞大学)
- Manning College of Information and Computer Science(曼宁信息与计算机科学学院)
- Dell AI Research Office(戴尔人工智能研究办公室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 H-UVFAs,将多目标通用价值函数逼近扩展到基于 options 的分层强化学习,并通过监督与强化学习方法学习分层嵌入,验证其泛化性能优于 UVFAs。
AI中文摘要:
在为强化学习价值函数的多目标集合构建通用逼近器方面,已经取得了关键进展;这些价值函数是以参数化方式估计状态长期回报的关键要素。我们通过引入分层通用价值函数逼近器(H-UVFAs),利用 options(选项)框架将其扩展到分层强化学习。这使我们能够利用时间抽象设置中预期的扩展、规划和泛化所带来的额外优势。我们开发了监督学习和强化学习方法,用于学习两个分层价值函数 $Q(s, g, o; θ)$ 与 $Q(s, g, o, a; θ)$ 中状态、目标、选项和动作的嵌入。最后,我们证明了 HUVFA 的泛化能力,并表明它们优于相应的 UVFA。
英文摘要:
There have been key advancements to building universal approximators for multi-goal collections of reinforcement learning value functions -- key elements in estimating long-term returns of states in a parameterized manner. We extend this to hierarchical reinforcement learning, using the options framework, by introducing hierarchical universal value function approximators (H-UVFAs). This allows us to leverage the added benefits of scaling, planning, and generalization expected in temporal abstraction settings. We develop supervised and reinforcement learning methods for learning embeddings of the states, goals, options, and actions in the two hierarchical value functions: $Q(s, g, o; θ)$ and $Q(s, g, o, a; θ)$. Finally we demonstrate generalization of the HUVFAs and show they outperform corresponding UVFAs.