arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03260quant-phcs.ARcs.SYeess.SY

共同编译:通过多编译实现高吞吐量分布式量子计算

Compiling Together: High-Throughput Distributed Quantum Computing via Multi-Compilation

Yipei Liu, Sen Zhang, Zebo Yang, Lei Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对分布式量子计算中纠缠资源稀缺限制吞吐量的问题,提出多编译方法,通过并发执行同一电路的多种等价编译并汇总样本,将空闲贝尔态转化为额外射击,显著提升吞吐量和保真度。

中文摘要 AI 辅助

量子计算是解决经典机器难以处理的问题的一种有前景的范式,但实现这一前景所需的量子比特数量远超单个处理器所能提供的。分布式量子计算(DQC)通过连接多个量子处理单元(QPU)进行扩展,但代价是纠缠成为稀缺资源:每个远程门操作消耗一对贝尔态,而QPU间链路生成贝尔态的速率比本地门操作慢几个数量级。由于量子程序被重复执行,这一速率限制了完成射击(shots)的速度,从而也限制了获取结果的速度。现有的DQC编译器为每个电路生成单一实现,因此吞吐量受限于最繁忙的链路,而其他链路则处于空闲状态。我们观察到,同一电路的不同编译在逻辑上是等价的,但会加重不同链路的负担;并发执行这些编译并汇总其样本,可以将空闲的贝尔态转化为额外的射击。我们将编译的联合选择和在每条链路贝尔态容量限制下的射击分配问题形式化为候选约束最大射击分配(CMA)问题,证明其NP难,并分别用动态规划及其近似变体AppDP、紧凑型混合整数线性规划(MILP)以及贪心启发式算法Effi求解。在六QPU网络的模拟中,多编译相比单一编译将吞吐量提高了2至4.5倍,并在相同贝尔态预算下相应提升了输出保真度。在硬件上,它将实测保真度提高了最多93%,并在最多快10倍的时间内达到单一编译的保真度。扩展到36个QPU和144量子比特电路时,MILP在48种配置上平均将吞吐量比单一编译提高了76.8%,而Effi在毫秒级时间内完成分配。

英文摘要

Quantum computing is a promising paradigm for problems that are challenging for classical machines, but realizing that promise requires far more qubits than a single processor can offer. Distributed quantum computing (DQC) scales out by connecting multiple quantum processing units (QPUs), at the cost of making entanglement the scarce resource: every remote gate consumes a Bell pair, and inter-QPU links generate Bell pairs at finite rates orders of magnitude slower than local gates. Since quantum programs are executed repeatedly, this rate bounds how fast shots complete and thus how fast results are obtained. Existing DQC compilers emit a single implementation per circuit, so throughput is capped by its busiest link while other links stay idle. We observe that alternative compilations of the same circuit are logically equivalent yet stress different links; executing them concurrently and pooling their samples converts idle Bell pairs into additional shots. We formulate the joint selection of compilations and allocation of shots under per-link Bell-pair capacities as Candidate-Constrained Max-Shot Allocation (CMA), prove it NP-hard, and solve it with a dynamic program and its approximate variant AppDP, a compact MILP, and a greedy heuristic Effi. In simulation on six-QPU networks, multi-compilation raises throughput by 2-4.5* over the single compilation and correspondingly improves output fidelity under equal Bell-pair budgets. On hardware, it increases measured fidelity by up to 93% and reaches the single-compilation fidelity in up to 10* less time. Scaling to 36 QPUs and 144-qubit circuits, MILP improves throughput over the single compilation by 76.8% on average across 48 configurations while Effi allocates in milliseconds.

发表机构

  • George Mason University(乔治梅森大学)
  • Florida Atlantic University(佛罗里达大西洋大学)

机构由 AI 辅助整理,请以论文原文为准。

↑