arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27947quant-phcs.DC

面向量子电路裁剪后处理重构的CPU+DCU异构并行框架

A CPU+DCU Heterogeneous Parallel Framework for Post-Processing Reconstruction in Quantum Circuit Cutting

Qingqing Jiang, Weidong Liu, Yufu Liu, Ruiqing He, Jiandong Shang, Hengliang Guo, Qiang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对量子电路裁剪后处理的计算与存储瓶颈,提出CPU+DCU异构并行框架,实现非零概率态重构,在嵩山超算上获最高259倍加速,可完成百量子比特规模任务。

中文摘要 AI 辅助

在含噪声中等规模量子(NISQ)时代,量子比特资源有限,难以直接在真实硬件上执行大规模量子电路。量子电路裁剪通过将大规模电路分解为多个小子电路缓解该限制,但会将大量开销转移到经典后处理中。随着电路规模、复杂度和裁剪数量增加,重构成为主要的计算和存储瓶颈。本文提出一种用于电路裁剪后处理重构的CPU+DCU异构并行框架,该框架未构建稠密的2ⁿ维概率向量,也未仅返回高概率态,而是从子电路测量结果中重构原始输出分布中的非零概率态。该框架结合了异构CPU+DCU执行、用于64位以上全局基态索引的高低字整数表示,以及跨越设备内存、主机内存和外存的三级协同存储机制。在嵩山超级计算机上的实验表明,该框架在保持高重构保真度的同时,针对线性簇态实现了比优化串行基准最高259倍的加速,针对随机电路实现了比同构CPU并行方法最高4倍的加速,还可完成百量子比特规模的重构任务。这些结果表明,面向高性能计算(HPC)的异构重构可有效缓解经典后处理瓶颈,提升重构可扩展性。

英文摘要

In the NISQ era, limited qubit resources make it difficult to execute large quantum circuits directly on real hardware. Quantum circuit cutting mitigates this limitation by decomposing a large circuit into smaller subcircuits, but it shifts substantial overhead to classical post-processing. As circuit size, complexity, and cut count increase, reconstruction becomes a major computational and storage bottleneck. This paper presents a CPU+DCU heterogeneous parallel framework for circuit-cutting post-processing reconstruction. Instead of constructing a dense $2^n$-dimensional probability vector or returning only high-probability states, the framework reconstructs the nonzero-probability states in the original output distribution from subcircuit measurement results. It combines heterogeneous CPU+DCU execution with a high/low-word integer representation for global basis-state indices beyond 64 bits and a three-level cooperative storage mechanism spanning device memory, host memory, and out-of-core storage. Experiments on the Songshan supercomputer show that the framework maintains high reconstruction fidelity while achieving up to $259\times$ speedup over an optimized serial baseline on linear-cluster states and up to $4\times$ speedup over a homogeneous CPU-parallel method on random circuits. The framework can also complete reconstruction tasks at the hundred-qubit scale. These results demonstrate that HPC-oriented heterogeneous reconstruction can effectively alleviate the classical post-processing bottleneck and improve reconstruction scalability.

↑