arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Aicir:支持昇腾NPU的全栈量子电路模拟器

Aicir: A Full-Stack Quantum Circuit Simulator with AscendNPU Support

Xian Lu, Xinying Li, Fei Wang, Shuai Hou, Chengkang Pan, Xin Yi, Yongmei Li

arXiv 2608.09733首次发表:更新:

AI 中文总结

本研究开发了支持昇腾NPU原生后端的全栈量子电路模拟器Aicir,其CPU性能与成熟模拟器相当,可在多NPU间分布式执行并保留反向微分,填补了原生支持NPU的量子电路模拟器缺口。

AI 中文摘要

量子计算是研究经典方法难以解决问题的极具前景的方式,但当前量子硬件仍面临规模、噪声和保真度的限制,在物理机器上运行量子算法成本也较高。因此,量子电路模拟器至关重要,它能让研究人员在使用量子硬件前,在经典计算机上设计和测试算法。大多数高性能模拟器提供GPU后端,而很少有原生支持NPU的,这一缺口限制了量子算法研究可用的计算平台。我们开发了Aicir,提供支持华为昇腾(Ascend)NPU原生后端的全栈量子电路模拟器,它通过一个编程模型连接电路构建、多种态表示、测量、微分、变分算法、量子机器学习和量子架构搜索,还支持噪声模拟、张量网络与矩阵乘积态引擎,以及分布式态模拟。在NPU上,配对实张量、固定秩门视图和硬件特定公式,使测试的模拟路径保留在设备上,相同表示让Aicir可在2^p个NPU间划分态,同时保留反向模式微分。我们在禁用CPU fallback的情况下验证了原生执行,并在2、4、8个NPU上检查了分布式通信和梯度。对于测试的融合分层电路,Aicir的CPU运行时间为Qiskit Aer的0.97至1.28倍,Cirq的0.76至1.10倍,其CPU执行与该工作负载的成熟模拟器处于同一范围,而NPU测试验证了正确的原生执行,而非CPU到NPU的加速。

英文摘要

Quantum computing is a promising way to study problems that are difficult for classical methods, but current quantum hardware still faces limits in scale, noise, and fidelity. Running quantum algorithms on physical machines can also be costly. Quantum circuit simulators therefore remain important because they let researchers design and test algorithms on classical computers before using quantum hardware. Most high-performance simulators provide GPU backends, while few offer native support for NPUs. This gap limits the computing platforms available for quantum-algorithm research. We developed Aicir to provide a full-stack quantum circuit simulator with a native Huawei Ascend NPU backend. Aicir connects circuit construction, several state representations, measurement, differentiation, variational algorithms, quantum machine learning, and quantum architecture search through one programming model. It also supports noise simulation, tensor-network and matrix-product-state engines, and distributed state simulation. On the NPU, paired real tensors, fixed-rank gate views, and hardware-specific formulas keep the tested simulation paths on the device. The same representation lets Aicir partition a state across $2^{p}$ NPUs while retaining reverse-mode differentiation. We validated native execution with CPU fallback disabled and checked distributed communication and gradients on 2, 4, and 8 NPUs. For the tested fused layered circuits, Aicir's CPU runtime is within $0.97$--$1.28\times$ that of Qiskit Aer and $0.76$--$1.10\times$ that of Cirq. These results place its CPU execution in the same range as established simulators for this workload, while the NPU tests establish correct native execution rather than CPU-to-NPU speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑