arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成模型下强化学习的改进量子算法

Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model

Joao F. Doriguello

arXiv 2608.02826首次发表:更新:

发表机构

HUN-REN Alfréd Rényi Institute of Mathematics(HUN-REN 阿尔弗雷德·雷尼数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对两类马尔可夫决策过程,提出结合值迭代与量子子例程的新型量子强化学习算法,其查询复杂度优于现有成果,接近量子下界。

AI 中文摘要

强化学习是机器学习的子领域,研究智能体如何与环境交互以获取尽可能大的奖励。研究此类交互的标准方法是通过马尔可夫决策过程(MDP),任务是选择最优策略——即告知智能体采取何种动作的函数。本研究探讨有限时间范围和无限时间范围折扣这两类MDP,提出用于计算近似最优策略的新型量子算法。该算法结合标准值迭代与量子均值估计、量子最大值查找等量子子例程,并加入样本最优经典算法的技术,所得查询复杂度优于现有研究,接近已确立的量子下界。

英文摘要

Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible. A standard approach to study such interaction is through Markov Decision Processes (MDPs) and the task of choosing an optimal policy --- a function that tells the agent which action to take. In this work, we study two types of MDPs --- finite-horizon and infinite-horizon discounted --- and propose new quantum algorithms for computing approximate optimal policies. Our quantum algorithms are based on a new combination of standard value iteration and quantum subroutines like quantum mean estimation and quantum maximum finding, overall enhanced with techniques from sample-optimal classical algorithms. Our resulting query complexities improve upon previous works, thus approaching already established quantum lower bounds.

Comments22 pages. Comments welcome

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑