AI 中文总结
研究针对混合薛定谔 - 费曼量子电路仿真实际运行时间受跨边界双量子比特门指数级路径增长主导的问题,提出HSF-S框架,通过特定方法抑制跨边界相互作用,降低路径成本,在基准电路上大幅提高可处理性并实现加速。
AI 中文摘要
混合薛定谔 - 费曼(HSF)仿真为精确量子电路仿真提供了有吸引力的内存 - 路径权衡,但实际运行时间常受跨边界双量子比特门指数级路径增长主导。现有GPU和FPGA量子模拟器大多针对全状态薛定谔执行进行优化,与HSF以路径为中心的工作流程不匹配。本文提出HSF-S,一个用于基于HSF的精确量子电路仿真的编译器 - 加速器协同设计框架。HSF-S将输入电路降低到与HSF兼容的基,制定秩感知有效路径成本模型,并应用保持依赖的重新排序以及折扣增益SWAP插入来抑制反复出现的跨边界相互作用,同时保持精确的电路语义。一个无回归选择器确保编译后的电路相对于朴素降低的基线不会增加有效路径成本。我们进一步设计了专用的HSF-S加速器和执行流程,并将它们集成到一个独立处理器中,以实现高效的每条路径双切片评估和最终累加,而无需实现完整的状态向量。在56个基准电路上,HSF-S将参考幅度匹配到浮点精度范围内,有效路径成本降低高达90.0%,并显著提高了实际可处理性,包括在1小时预算下将代表性的超时时间减少到亚秒级。在生成的编译工作负载上,HSF-S处理器原型实现了高达4.34倍的额外加速。
英文摘要
Hybrid Schrodinger-Feynman (HSF) simulation offers an attractive memory-path tradeoff for exact quantum-circuit emulation, but its practical runtime is often dominated by exponential path growth from cross-boundary two-qubit gates. Existing GPU and FPGA quantum simulators are largely optimized for full-state Schrodinger execution and therefore do not align well with HSF's path-centric workflow. This paper presents HSF-S, a compiler-accelerator co-designed framework for exact HSF-based quantum circuit emulation. HSF-S lowers input circuits to an HSF-compatible basis, formulates a rank-aware effective path-cost model, and applies dependency-preserving reordering together with discounted-gain SWAP insertion to suppress recurring cross-boundary interactions while preserving exact circuit semantics. A regression-free selector guarantees that the compiled circuit never increases effective path cost relative to the naive lowered baseline. We further design a dedicated HSF-S accelerator and execution flow, and integrate them into a stand-alone processor for efficient per-path dual-slice evaluation and final accumulation without materializing the full state vector. Across 56 benchmark circuits, HSF-S matches reference amplitudes to within floating-point precision, reduces effective path cost by up to 90.0%, and substantially improves practical tractability, including representative timeout-to-sub-second reductions under a 1-hour budget. On the resulting compiled workloads, the HSF-S processor prototype delivers up to 4.34x additional speedup.
Comments11 pages, 8 figures, 2 tables. This paper is accepted for ICCAD 2026