AI 中文总结
研究混合量子-经典程序中无贡献测量结果的检测问题,提出语义感知主机端静态分析方法,经实验验证该方法能有效识别并移除更多此类结果,还通过降低程序表示和实现CUDA后端实现加速,提升了分析效率。
AI 中文摘要
混合程序将量子电路与消耗测量结果的经典主机程序相结合。在这类程序中,某些结果可能在语法上被主机读取,但在语义上并无贡献:改变结果不会改变返回值。这些结果会掩盖仅相对于主机语义而言无效的门,因此电路局部优化器无法察觉。我们提出一种语义感知的主机端静态分析方法,通过抽象解释识别无贡献的测量结果,并证明其正确性。我们实现了该分析并在24个跨量子化学、优化、量子机器学习和量子金融的忠实应用混合工作负载上进行评估。与语法活跃性基线相比,我们的分析识别出的无贡献测量数量多出4倍以上,平均能单独移除37.98%的总门。即使在Qiskit、t|ket⟩和PyZX等先进优化器对电路进行优化后,我们的分析仍能移除超过30%的优化后门,表明我们分析所揭示的主机语义机会未被电路局部优化涵盖。为扩展我们的分析,我们进一步将主机程序降低到SSA风格的分层中间表示,以暴露用于GPU执行的分层并行性,并实现了CUDA后端。我们证明这种降低保留了分析结果,评估显示随着结构并行性增加,速度比顺序基线提高了6.53倍。
英文摘要
Hybrid programs combine a quantum circuit with a classical host program that consumes measurement outcomes. In such programs, an outcome may be syntactically read by the host but semantically non-contributory: changing the outcome cannot change the returned value. Such outcomes obscure gates that are dead only relative to the host semantics, and are therefore invisible to circuit-local optimizers. We present a semantics-aware host-side static analysis that identifies non-contributory measurement outcomes by abstract interpretation, and prove its soundness. We implement the analysis and evaluate it on $24$ application-faithful hybrid workloads across quantum chemistry, optimization, quantum machine learning, and quantum finance. Compared with a syntactic liveness baseline, our analysis identifies more than $4\times$ as many non-contributory measurements, and it standalone enables the removal of $37.98\%$ of total gates on average. Even after the state-of-the-art optimizers like Qiskit, t|ket$\rangle$, and PyZX have already optimized the circuits, our analysis still enables removal of more than $30\%$ of the post-optimized gates, showing that the host-semantic opportunities exposed by our analysis are not subsumed by circuit-local optimization. To scale our analysis, we further lower host programs to an SSA-style levelized intermediate representation that exposes level-wise parallelism for GPU execution, and implement a CUDA backend. We prove that this lowering preserves the analysis result, and the evaluation shows speedups of up to $6.53\times$ over a sequential baseline as structural parallelism increases.
CommentsAccepted at 33rd Static Analysis Symposium (SAS 2026), https://conf.researchr.org/home/splash-issta-2026/sas-2026. This is the full version that includes all proofs and technical details in the appendices, which are omitted in the conference manuscript for conciseness