arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03987quant-ph

实值化张量网络:在实值矩阵加速器上的量子电路模拟

Realified tensor networks: quantum circuit simulation on real-valued matrix accelerators

Yusheng Zhao, Xiwei Pan, Enji Xiong, Chengkai Zhu, Jinguo Liu

AI总结:

该研究提出实值化张量网络方法,将复值量子模拟张量网络映射为实值网络,在昇腾910 NPU等实值矩阵加速器上实现高效量子电路模拟,性能优于基线方法。

AI中文摘要:

张量网络收缩可用于模拟量子电路,但现代矩阵加速器(如NPU、TPU)仅提供实值通用矩阵乘法(GEMM)流水线,因此量子模拟所需的复值网络必须在软件中重构。我们通过实值化重写解决这一不匹配问题,将任意复值张量网络映射为实值张量网络。在两个复值张量的每次合并中,一个秩-3结构张量实现高斯三乘法(3M)公式;与一个或无复值操作数的收缩仅需两次或一次实值乘积。我们证明了一个紧成本定律:实值乘法的开销为1 + 2m + r,其中m和r分别是双复值操作数收缩和单复值操作数收缩的体积分数,相对于实值收缩的开销绝不超过3倍,且所有中间张量的大小最多加倍。在67个电路(随机电路、Clifford+$T$电路、QAOA、VQE)上,该定律在实值到复值的所有范围以及复值门的位置、数量上均成立,成本由收缩顺序而非其他因素决定。在67个电路中的66个上,从复值网络转移而来的收缩顺序的相对算术成本差距低于5×10⁻⁴;例外情况在少量低温模拟退火步骤后即可消除。在昇腾910 NPU上,该重写在全部12个随机电路和55个结构化单元中的52个上,均优于四倍实值GEMM基线和每个GEMM的高斯降阶方法(三个单元的速度慢不超过12%);四倍GEMM基线的中位数速度比分别慢1.7倍(随机电路)和1.4倍(结构化单元)。实值化使复值张量网络收缩可在仅支持实值的矩阵引擎上原生运行。

英文摘要:

Tensor-network contraction simulates quantum circuits, but modern matrix accelerators (NPUs, TPUs) expose only real GEMM pipelines, so the complex networks of quantum simulation must be reconstructed in software. We resolve the mismatch by a realification rewrite that maps any complex tensor network to a real one. At each merge of two complex tensors, a rank-3 structure tensor realizes Gauss's three-multiplication (3M) formula; contractions with one or no complex operand need only two or one real products. We prove a tight cost law: overhead $1 + 2m + r$ in real multiplications, where $m$ and $r$ are the volume fractions of two- and one-complex-operand contractions, never exceeding $3\times$ relative to real contraction, with every intermediate at most doubled in size. On 67 circuits (random, Clifford+$T$, QAOA, VQE), the law holds across the real-to-complex range and complex-gate placement, not count, governs cost. Contraction orders transfer from the complex network with a relative arithmetic-cost gap below $5\times 10^{-4}$ on 66 of 67 circuits; the exception closes under a few steps of low-temperature simulated annealing. On an Ascend 910 NPU the rewrite beat both the four-real-GEMM baseline and a per-GEMM Gauss lowering on all twelve random circuits and on 52 of 55 structured cells (three cells slower by at most 12\%); the four-GEMM baseline was slower by a median $1.7\times$ (random) and $1.4\times$ (structured). Realification makes complex tensor-network contraction native to real-only matrix engines.

↑