AI 中文总结
研究针对 GPU 噪声量子电路模拟,实现方差降低轨迹展开(投影器和模拟采样)并验证。在 GPU 上投影器展开比 Qiskit-Aer 方法所需轨迹大幅减少,明确两种展开适用噪声范围,指出 Qiskit-Aer 集成问题并给出解锁技术的最小更改。
AI 中文摘要
蒙特卡罗轨迹(量子跳跃)方法是模拟噪声量子电路的实用途径,一旦精确密度矩阵方法因其\(4^n\)的内存成本而不可行。其瓶颈在于估计器方差,解决一个期望值可能需要数千条轨迹。近期张量网络工作表明,方差降低展开——投影器和模拟采样——能大幅降低方差,但仅适用于 CPU 矩阵乘积态后端,无法进入生产工具。我们在 GPU 密集态矢量轨迹引擎上实现了这两种展开,并与精确密度矩阵进行验证(理想电路保真度\(1 - 2.×10^{-16}\);\(1/\sqrt{N}\)收敛;所有展开在迹距离\(<0.01\)时无偏)。在单个消费级 GPU 上,投影器展开在\(n = 10\)时达到目标标准误差所需轨迹比 Qiskit-Aer 的 \texttt{batched\_shots\_gpu} 少\(20.8\)倍,在\(n = 8\)至\(20\)时为\(19\)至\(26\)倍。一个机制图表明模拟采样在弱噪声下最佳,投影器在强噪声下最佳,交叉点接近\(\gamma t\approx0.35\)。我们还报告了一个系统发现:Qiskit-Aer 在通道级别应用噪声并在应用时重建规范的克劳斯分解,丢弃任何用户提供的展开,因此方差降低展开无法通过其公共 API 提供。由于 Aer 的玻恩法则坍缩机制已经存在,我们指定了一个最小更改以在生产中解锁该技术。
英文摘要
Monte-Carlo trajectory (quantum-jump) methods are the practical route to simulating noisy quantum circuits once the exact density-matrix method is precluded by its $4^n$ memory cost. Their bottleneck is estimator variance: resolving one expectation value can demand thousands of trajectories. Recent tensor-network work shows that \emph{variance-reduced unravelings} -- projector and analog sampling -- sharply cut this variance, but only on CPU matrix-product-state backends, with no path into production tooling. We implement both unravelings on a \emph{GPU dense-statevector} trajectory engine and validate them against the exact density matrix (ideal-circuit fidelity $1-2.2\times10^{-16}$; $1/\sqrt{N}$ convergence; all unravelings unbiased to trace distance $<0.01$). On a single consumer GPU, projector unraveling reaches a target standard error with $20.8\times$ fewer trajectories than Qiskit-Aer's \texttt{batched\_shots\_gpu} at $n=10$, a factor that holds at $19$--$26\times$ across $n=8$--$20$. A regime map places analog sampling optimal at weak noise and projector at strong noise, crossing near $γt\approx0.35$. We further report a systems finding: Qiskit-Aer applies noise at the \emph{channel} level and reconstructs a canonical Kraus decomposition at apply time, discarding any user-supplied unraveling, so variance-reduced unravelings cannot be delivered through its public API. Because Aer's Born-rule collapse machinery already exists, we specify a minimal change that would unlock the technique in production.
Comments20 pages, 6 figures, 4 tables. Submitted for publication