发表机构
Stony Brook University; Virginia Tech(石溪大学; 弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用强化学习(DQN+MPNN)为量子网络中的同时纠缠请求制定调度策略,在降低链路激活概率下保持高成功率,并借助LLM提取可解释的启发式策略。
AI 中文摘要
未来的量子网络将利用纠缠来执行众多任务,例如远距离传输量子信息、分布式量子计算和量子传感。通常,这些任务需要在网络的不同区域同时执行,同时最小化资源和延迟。因此,我们将需要用于调度链路级纠缠资源的策略,并利用链路级纠缠为每项任务创建所需的各种形式的多部分纠缠。在这项工作中,我们使用强化学习来解决这个问题。我们为该问题制定了一个马尔可夫决策过程,并使用带有消息传递神经网络(MPNN)的双深度Q网络(DQN)、经验回放缓冲区和课程训练来获得策略。关键的物理参数是链路级纠缠生成的概率,即链路激活概率。我们表明,对于一组物理相关的网络拓扑,我们的策略在链路激活概率比基线启发式方法低高达71%的情况下,仍能保持100%的成功率。然后,我们检查了一个额外的约束,即实验(任务)放置被限制在特定的硬件类型上,并展示了在性能上相对于启发式方法的类似优势,我们的策略在链路激活概率降低高达59%的情况下,仍能保持至少80%的成功率。最后,我们探索了通过定义指标来解释所学策略的方法,这些指标能够对模型的行为得出结论,并通过让大型语言模型(LLM)根据DQN训练策略所采取的示例动作来推导出一种新颖的启发式方法。我们发现,LLM启发式方法在性能上与DQN训练策略相似,这表明对于直接训练变得计算昂贵的的大型量子网络,这是一种有前景的可解释策略提取方法。
英文摘要
Future quantum networks will make use of entanglement to perform numerous tasks, such as sending quantum information over long distances, distributed quantum computing, and quantum sensing. In general, these tasks will need to be performed simultaneously in various regions of a network, while minimizing resources and latency. We will thus require policies for scheduling link-level entanglement resources, and using the link-level entanglement to create various forms of multipartite entanglement required for every task. In this work, we address this problem using reinforcement learning. We formulate a Markov Decision Process for the problem and use double deep Q-networks (DQN) with Message Passing Neural Networks (MPNNs), experience replay buffers, and curriculum training to obtain policies. The key physical parameter is the probability of link-level entanglement generation, i.e., the link activation probability. We show that our policies maintain 100% success for up to 71% lower link activation probability than the baseline heuristics for a set of physically relevant network topologies. We then examine an additional constraint where experiment (task) placements are restricted to specific hardware types and demonstrate a similar advantage in performance over heuristics, with our policy maintaining at least an 80% success rate for up to a 59% lower link activation probability. Finally, we explore methods to interpret the learned policy by defining metrics enabling conclusions to be drawn about the model's behavior and by tasking a large language model (LLM) to derive a novel heuristic given example actions taken by the DQN-trained policy. We find that the LLM heuristic performs similarly to the DQN-trained policy in performance, indicating a promising method for interpretable policy extraction for large quantum networks, where direct training becomes computationally expensive.