arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向通信高效分布式量子电路编译的架构感知强化学习

Architecture-Aware Reinforcement Learning for Communication-Efficient Distributed Quantum Circuit Compilation

Chien-Tung Kuo, Felix Burt, Samuel Yen-Chi Chen, Kin K. Leung, Kuan-Cheng Chen

arXiv 2608.06892首次发表:更新:

AI 中文总结

该研究针对分布式量子计算的通信感知编译问题,提出架构感知强化学习框架,经实验验证其在结构化工作负载上性能与顶尖启发式算法相当,为手动启发式算法提供了灵活替代方案。

AI 中文摘要

分布式量子计算为执行超出单个量子处理单元(QPU)容量限制的量子电路提供了可扩展途径,但它引入了涉及严格硬件约束和电路依赖关系的通信感知编译问题。本文提出一种架构感知强化学习框架,将分布式量子编译表述为约束马尔可夫决策过程(MDP)。编译器级通信操作会动态更新逻辑量子比特的布局,以实现后续门操作的执行。异构图模型用于表征硬件、逻辑量子比特和电路操作之间的交互,而通过近端策略优化(Proximal Policy Optimization)训练的策略则可优化EPR对的消耗和通信总时长。对基准电路的评估表明,该策略在结构化工作负载上与最先进的启发式算法表现相当,而前瞻奖励塑形则在非结构化电路上带来了适度的性能提升。这些结果证明,强化学习是手动启发式算法的灵活替代方案,不过可扩展性仍是其实际应用的关键瓶颈。

英文摘要

Distributed quantum computing provides a scalable route for executing quantum circuits beyond the capacity limits of a single quantum processing unit (QPU), but it introduces a communication-aware compilation problem involving strict hardware constraints and circuit dependencies. This paper presents an architecture-aware reinforcement-learning framework that formulates distributed quantum compilation as a constrained Markov Decision Process (MDP). The compiler-level communication actions dynamically update logical-qubit placement and enable subsequent gate execution. A heterogeneous graph model represents interactions among hardware, logical qubits, and circuit operations, while a policy trained via Proximal Policy Optimization optimizes EPR-pair consumption and communication makespan. Evaluation across benchmark circuits shows that our policy matches state-of-the-art heuristics on structured workloads, with lookahead reward shaping yielding modest improvements on unstructured circuits. These results demonstrate that reinforcement learning is a flexible alternative to manual heuristics, though scalability remains a key bottleneck for practical use.

Comments11 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑