发表机构
HUN-REN Alfréd Rényi Institute of Mathematics(HUN-REN 阿尔弗雷德·雷尼数学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对两类马尔可夫决策过程,提出结合值迭代与量子子例程的新型量子强化学习算法,其查询复杂度优于现有成果,接近量子下界。
AI 中文摘要
强化学习是机器学习的子领域,研究智能体如何与环境交互以获取尽可能大的奖励。研究此类交互的标准方法是通过马尔可夫决策过程(MDP),任务是选择最优策略——即告知智能体采取何种动作的函数。本研究探讨有限时间范围和无限时间范围折扣这两类MDP,提出用于计算近似最优策略的新型量子算法。该算法结合标准值迭代与量子均值估计、量子最大值查找等量子子例程,并加入样本最优经典算法的技术,所得查询复杂度优于现有研究,接近已确立的量子下界。
英文摘要
Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible. A standard approach to study such interaction is through Markov Decision Processes (MDPs) and the task of choosing an optimal policy --- a function that tells the agent which action to take. In this work, we study two types of MDPs --- finite-horizon and infinite-horizon discounted --- and propose new quantum algorithms for computing approximate optimal policies. Our quantum algorithms are based on a new combination of standard value iteration and quantum subroutines like quantum mean estimation and quantum maximum finding, overall enhanced with techniques from sample-optimal classical algorithms. Our resulting query complexities improve upon previous works, thus approaching already established quantum lower bounds.
Comments22 pages. Comments welcome