arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12172quant-ph

可扩展量子计算的并行电路执行

Parallel Circuit Execution for Scalable Quantum Computation

Avimita Chatterjee, W. Michael Brown, Siyuan Niu, Wibe Albert de Jong, Thomas Lubinski

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种错误与拓扑感知的并行电路执行方法,将多个独立电路映射到大型QPU的不同区域,在IBM 156量子比特处理器上实现3.5-5.5倍加速并保留高保真度,在GPU模拟中达13.8倍加速,显著降低量子计算执行成本。

中文摘要 AI 辅助

当今的量子处理器拥有数十到数百个物理量子比特,但在各种硬件平台上,可靠执行任意电路仍局限于少于30个纠缠量子比特。在先前工作的基础上,我们引入了一种错误感知和拓扑感知的方法,将多个独立电路映射到单个大规模QPU的不相交区域,以实现并行电路执行。对于许多规模相似的电路应用,如哈密顿量模拟中的可观测量估计,这种方法可以减少计费的QPU执行时间,理想加速比与可用分区的数量成正比。我们在IBM的156量子比特ibm_boston处理器上,使用标准的QED-C基准和基于哈密顿量的可观测量估计工作负载演示了该方法。与标准顺序执行相比,并行执行将计费执行时间减少了3.5-5.5倍,同时保留了83-92%的顺序保真度。我们进一步使用GPU加速的经典模拟评估了并行电路执行的扩展性,通过MPI将测量电路分布到多个GPU上。使用NERSC Perlmutter系统上的CUDA-Q,我们在16个GPU上实现了高达13.8倍的加速(86%的并行效率),用于H2电子结构模拟,并在多个哈密顿量和电路数量上评估了扩展性。这些结果提供了未来量子硬件上并行执行可能最终接近的性能上限的指示。两种执行模式都作为QED-C面向应用基准套件中的运行时选项实现。总之,结果表明电路级并行可以降低当前量子硬件上的执行成本和GPU集群上的模拟时间,并且随着设备质量和量子比特数量的增加,有可能带来更大的收益。

英文摘要

Today's quantum processors have tens to hundreds of physical qubits, but reliable execution of arbitrary circuits remains limited to fewer than 30 entangled qubits across hardware modalities. Building upon prior work, we introduce an error- and topology-aware method for mapping multiple independent circuits onto disjoint regions of a single large-scale QPU for parallel circuit execution. For applications with many similarly sized circuits, such as observable estimation for Hamiltonian simulation, this approach can reduce billed QPU execution time, with ideal speedup proportional to the number of usable partitions. We demonstrate the approach on IBM's 156-qubit ibm_boston processor using standard QED-C benchmark and Hamiltonian-based observable-estimation workloads. Compared with standard sequential execution, parallel execution reduces billed execution time by 3.5-5.5x while retaining 83-92% of the sequential fidelity. We further evaluate the scaling of parallel circuit execution using GPU-accelerated classical simulation, distributing measurement circuits across GPUs via MPI. Using CUDA-Q on the NERSC Perlmutter system, we achieve up to 13.8x speedup on 16 GPUs (86% parallel efficiency) for an H2 electronic-structure simulation, with scaling evaluated across multiple Hamiltonians and circuit counts. These results provide an indication of the performance ceiling that parallel execution on future quantum hardware may eventually approach. Both execution modes are implemented as a runtime option within the QED-C Application-Oriented Benchmark suite. Together, the results show that circuit-level parallelism can reduce execution cost on current quantum hardware and simulation time on GPU clusters, with the potential for greater benefits as device quality and qubit counts increase.

发表机构

  • Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室)
  • NVIDIA(英伟达)
  • University of Central Florida(中佛罗里达大学)
  • Quantum Computing Data(量子计算数据)
  • Cascade Quantum(瀑布量子)
  • QED-C Technical Advisory Committee – Standards(QED-C技术咨询委员会-标准)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑