迈向基于代理的量子强化学习去量子化
Towards Surrogate Based Dequantization of Quantum Reinforcement Learning
- University of the Basque Country UPV/EHU(巴斯克大学 UPV/EHU)
- TECNALIA, Basque Research and Technology Alliance (BRTA)(TECNALIA,巴斯克研究与技术联盟(BRTA))
- Freie Universität Berlin(柏林自由大学)
- Helmholtz-Zentrum Berlin für Materialien und Energie(亥姆霍兹柏林材料与能源中心)
- IKERBASQUE, Basque Foundation for Science(IKERBASQUE,巴斯克科学基金会)
- Basque Center for Applied Mathematics (BCAM)(巴斯克应用数学中心(BCAM))
- African Institute for Mathematical Sciences (AIMS)(非洲数学科学研究所(AIMS))
- Stellenbosch University(斯泰伦博斯大学)
- National Institute for Theoretical and Computational Sciences (NITheCS)(国家理论与计算科学研究所(NITheCS))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过核化拟合Q迭代为量子Q学习提供去量子化保证,在均匀生成模型下给出有限样本分析,并给出充分条件以证明经典算法可匹配量子性能。
AI中文摘要:
近年来,参数化量子电路作为函数逼近器的实用性已被广泛研究。在强化学习的背景下,这种方法催生了诸如量子Q学习之类的变分量子算法。尽管这些方法在经验上显示出有前景的结果,并且可以为人为设计的问题提供可证明的优势,但对于实际相关问题,它们能否相对于经典方法提供可证明的量子优势仍不清楚。研究此问题的一种自然途径是通过去量子化的视角:即构建能够匹配量子变分方法性能的高效经典算法。基于近期针对监督学习的基于核的去量子化结果,我们朝着将该基于代理的去量子化方案扩展到强化学习迈出了步伐。具体而言,我们研究了具有均匀生成模型的简化强化学习设置,在该设置中可获得均匀随机的状态-动作样本,这模拟了在充分探索后从大型经验回放缓冲区中采样的场景。在此设置下,我们为经典核化拟合Q迭代提供了有限样本保证,其中经典核被设计为匹配特定参数化量子电路的归纳偏置。利用这些结果,我们随后提供了一组充分条件,涉及参数化量子电路的数据编码策略、相应的经典核以及问题结构,在这些条件下,核化拟合Q迭代在此简化设置中提供了对量子Q学习的有意义的去量子化。除了在这些条件满足时提供严格的去量子化保证外,这些结果还激励了在无法验证这些充分条件时,将核化拟合Q迭代用作去量子化启发式方法。
英文摘要:
In recent years, the utility of parameterized quantum circuits as function approximators has been widely studied. In the context of reinforcement learning, this approach has led to variational quantum algorithms such as quantum Q-learning. While these methods show promising empirical results, and can provide provable advantages for artificial problems, it remains unclear whether they can provide a provable quantum advantage over classical approaches for problems of practical relevance. A natural way to investigate this question is through the lens of dequantization: The construction of efficient classical algorithms capable of matching the performance of quantum variational methods. Building on recent kernel-based dequantization results for supervised learning, we take steps towards extending this surrogate-based dequantization program to reinforcement learning. Specifically, we study the simplified setting of reinforcement learning with a uniform generative model in which uniformly random state-action samples are available, which models the regime of sampling from a large experience replay buffer after sufficient exploration. Within this setting, we provide finite sample guarantees for classical kernelized Fitted Q-Iteration, with classical kernels designed to match the inductive bias of particular parameterized quantum circuits. Using these results, we then provide a set of sufficient conditions, on the data-encoding strategy of a parameterized quantum circuit, the corresponding classical kernel, and the problem structure, under which kernelized Fitted Q-Iteration provides a meaningful dequantization of quantum Q-learning, in this simplified setting. Apart from providing rigorous dequantization guarantees when these conditions are met, these results also motivate the use of kernelized fitted Q-iteration as a dequantization heuristic when these sufficient conditions cannot be verified.