量子人工智能会以路径为梦吗?通过格罗弗干涉的路径积分慢思考
Do Quantum AIs Dream in Paths? Path-Integral Slow Thinking through Grover Interference
浏览论文内容
中文总结 AI 辅助
本研究提出基于格罗弗干涉的路径积分慢思考量子AI方法,通过内部训练目标避免策略崩溃,在滑动拼图任务中优于经典对照,保留探索多样性并提升验证性能。
中文摘要 AI 辅助
基于可验证奖励的强化学习使大型语言模型能够进行慢思考,但同样的训练可能引发策略崩溃:概率集中于少数成功轨迹,探索性多样性被侵蚀。我们探究量子人工智能能否以不同方式实现慢思考。我们将慢思考表述为推理轨迹上的相干动力学,即一个离散路径积分,其中动作序列以叠加态共存,并在测量前重新组合。在我们的可训练实现中,一个精确验证器将整体划分为集体接受和拒绝的组件,这些组件在格罗弗振幅放大下发生干涉。当预放大成功概率位于一个解析确定的低于1的值时,有限的格罗弗演化达到最大化,因此推理本身定义了一个内部训练目标,消除了向单位成功率的单调压力。在2x3滑动拼图的精确态矢量模拟中,格罗弗训练在一轮中于32个问题的训练集上达到0.95的准确率,而最强经典对照为0.73。在保留问题上,专业化有代价:通过相同放大读出的未训练均匀策略在此解密集基准上仍是最强参考,且量子训练比经典训练保留多得多的保留准确率——在四轮匹配电路应用下,两个量子模型达到最强经典对照的3.2倍和3.9倍。在每个放大预算下,固定规模策略支持的训练问题数量也随预算增长快于匹配的经典重复。这些结果确立了基于格罗弗的路径积分慢思考实现:内部目标保留探索路径多样性,整体干涉将其转化为验证性能。
英文摘要
Reinforcement learning with verifiable rewards enables large language models to think slowly, but the same training can induce policy collapse: probability concentrates onto a few successful trajectories and exploratory diversity erodes. We ask whether quantum AI can realize slow thinking differently. We formulate slow thinking as coherent dynamics over reasoning trajectories, a discrete path integral in which action sequences coexist in superposition and recombine before measurement. In our trainable realization, an exact verifier partitions the ensemble into collective accepted and rejected components that interfere under Grover amplitude amplification. A finite Grover evolution is maximized when the pre-amplification success probability lies at an analytically determined value below one, so inference itself defines an interior training target and removes the monotonic pressure toward unit success. In exact statevector simulations of a 2x3 sliding puzzle, Grover training reaches accuracy 0.95 on a 32-question training set at one round, against 0.73 for the strongest classical control. On held-out questions specialization has a cost: an untrained uniform policy read out through the same amplification remains the strongest reference on this solution-dense benchmark, and quantum training preserves far more held-out accuracy than classical training - at four rounds with matched circuit applications the two quantum models reach 3.2 and 3.9 times the strongest classical controls. The number of training questions supported by fixed-size policies trained at each amplification budget also grows faster with the budget than with matched classical repetition. These results establish a Grover-based realization of path-integral slow thinking: the interior target preserves exploratory path diversity, and ensemble-level interference converts it into verified performance.
发表机构
- Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所)
- Shenzhen International Quantum Academy(深圳国际量子研究院)
- Hefei National Laboratory(合肥国家实验室)
机构由 AI 辅助整理,请以论文原文为准。