发表机构
BlueQubit Inc; Carnegie Mellon University(BlueQubit公司; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对GPU加速的MPS和PPS量子模拟方法进行基准测试,比较多个工具,发现MPS中GPU运行时随键维度次二次缩放,PPS中GPU在特定基准测试下加速显著且能达高精度,还提供模拟器特征描述及代码配置仓库。
AI 中文摘要
从业者越来越依赖托管模拟环境,但其性能特征记录不佳。我们对两种广泛使用的方法(矩阵乘积态(MPS)和泡利路径模拟(PPS))进行了GPU加速近似量子模拟的系统基准测试研究,将BlueQubit与AWS Braket、Quantum Rings、PPS-Qiskit等进行比较。对于MPS,GPU运行时随键维度呈次二次缩放,规模越大比CPU优势越明显。对于IBM 127量子比特踢伊辛基准测试的泡利路径模拟(PPS),在精细截断阈值下GPU加速高达1400倍,且是唯一能达到低于δ = 10⁻⁵精度的后端。我们还提供了这些模拟器在不同情况下的可重复特征描述,并提供了包含所有基准测试代码和配置的公共GitHub仓库。
英文摘要
Practitioners increasingly rely on hosted simulation environments, but their performance characteristics remain poorly documented. We present a systematic benchmarking study of GPU-accelerated approximate quantum simulation across two widely used methods: matrix product states (MPS) and Pauli path simulation (PPS), comparing BlueQubit (a hosted tool that handles hardware provisioning, simulator configuration, and job orchestration) against AWS Braket, Quantum Rings, Qiskit pauli-prop, and PauliPropagation (written in Julia). For MPS, we find that GPU runtime yields sub-quadratic scaling with bond dimension, with a growing advantage over CPU at increasing scale. For Pauli path simulation on IBM's 127-qubit kicked Ising benchmark, GPUs deliver up to ${\sim}1{,}700\times$ speedup at fine truncation thresholds ($δ= 2.5 \times 10^{-5}$, 27.6M Pauli terms), and are the only backends that reach accuracy regimes below $δ= 10^{-5}$, which remained inaccessible to the commodity CPU-based implementations and self-contained SDKs evaluated here. We also provide a reproducible characterization of these simulators across regimes, including tradeoffs that isolated evaluations do not show. All benchmarking code and configurations are in a public GitHub repository.
Comments12 pages, 12 figures