arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更少历史,更快路径:基于历史缩减、检查点和剪枝的分布式量子电路费曼模拟

Fewer Histories, Faster Paths: Distributed Quantum Circuit Feynman Simulation via History Reduction, Checkpointing, and Pruning

Frej Larssen, Luca Pennati, Erik M. Åsgrim, Ivy Peng, Stefano Markidis

arXiv 2608.01467首次发表:更新:

AI 中文总结

该研究提出分布式量子电路费曼模拟方法,通过历史缩减等技术解决路径和指数增长问题,在多类量子电路上实现高效精确模拟,获高并行效率。

AI 中文摘要

我们提出一种基于纯费曼历史求和公式的分布式精确稀疏输出量子电路模拟方法。该方法精确计算选定计算基幅,通过基于内部线分配、确定性传播、人工源、剪枝和检查点复用的缩减历史公式,解决路径和的指数增长问题。边界约束经确定性和保线门传播,仅在剩余歧义处引入显式分支变量;通过自动调谐的检查点分区捕获相关历史间的共享工作。并行执行模型结合请求输出的分解与并发历史评估,动态服务器-工作者架构缓解不规则分支和剪枝导致的负载不平衡。在所研究的电路族中,该方法适配缩减历史空间的不同结构 regime:QFT 反向分析下零个人工源,振幅放大中检查点和自动调谐带来显著加速,QAOA 中阈值剪枝实现运行时-保真度权衡;在量子行走电路上,其重构高达 100 量子比特的精确选定输出分布,在超级计算机 8192 个 CPU 核上达到 85% 的并行效率。

英文摘要

We present a distributed method for exact sparse-output quantum circuit simulation based on the pure Feynman sum-over-histories formulation. The method computes selected computational-basis amplitudes exactly and addresses the exponential growth of the path sum through a reduced history formulation based on internal-wire assignments, determinism propagation, artificial sources, pruning, and checkpointed reuse. Boundary constraints are propagated through deterministic and wire-preserving gates, and explicit branching variables are introduced only where residual ambiguity remains. Shared work across related histories is captured via an autotuned checkpointed partition. The parallel execution model combines decomposition over requested outputs with concurrent history evaluation, while a dynamic server-worker architecture mitigates load imbalance from irregular branching and pruning. Across the circuit families studied, the method adapts to different structural regimes of the reduced history space: zero artificial sources for QFT under backward analysis, substantial speedups from checkpointing and autotuning for amplitude amplification, and a runtime-fidelity tradeoff from threshold pruning for QAOA. On quantum walk circuits, it reconstructs exact selected-output distributions up to 100 qubits and achieves 85% parallel efficiency on 8,192 CPU cores of a supercomputer.

Comments11 pages, 12 figures. Accepted manuscript to appear in the proceedings of the IEEE International Conference on Quantum Computing and Engineering (QCE 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑