arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

输出等价性预言机遗漏了什么:量子转译器(Qiskit、tket、Cirq)中等价性不可见缺陷修复的实证研究

What Output-Equivalence Oracles Miss: An Empirical Study of Equivalence-Invisible Bug Fixes in Quantum Transpilers

Furqan Nasir, Arif Shah, Iftikhar Alam

arXiv 2609.13839首次发表:更新:

AI 中文总结

本研究实证测量量子转译器中输出等价性预言机遗漏的缺陷修复比例,发现Qiskit约28%的修复涉及等价性不可见故障,且该现象在tket和Cirq中复现,揭示了现有验证标准的盲区。

AI 中文摘要

量子编译器通过输出等价性预言机来评判正确性:编译后的电路必须计算与原始电路相同的酉算子,模去全局相位和量子比特布局置换。该预言机在构造上只检查该语义映射,而不检查电路自身的布局、置换或相位记录:那里的缺陷,或固定种子运行中的确定性问题,即使记录是公开的,也可能在未被察觉的情况下通过。我们衡量了在真实合并的编译器修复中这种情况发生的频率:一个系统识别的Qiskit转译器缺陷修复语料库,通过独立双编码、源码验证的表现分类法进行分类。68个修复中有19个(28%,95% Wilson置信区间19-40%)修复了该等价性筛选未捕获的故障,即使增加了编译有效性、电路质量和性能检查也是如此。一个次要的更宽窗口稳健性样本产生了几乎相同的比例(29/104,27.9%,置信区间20-37%)。即使仅限制在一个无条件等价性不可见通道(损坏的布局或置换记录),保守下限仍然存在:68个修复中有10个(15%,置信区间8-25%)低于28%的标题数字。该差距并非Qiskit特有:它在tket中重复出现(21个中有7个,33%),Cirq较小但一致。我们在五个廉价的PR级特征上未检测到系统性差异(19对49,功效不足无法排除适度的真实差异)。该类在不可见集合中占主导地位,集中在表示边界交叉处。我们发布了语料库、编码手册和编码工件。这里我们仅测量它。

英文摘要

Quantum compilers face the test oracle problem, judged by an output-equivalence oracle: the compiled circuit must compute the same unitary as the original, modulo global phase and a qubit-layout permutation. This oracle, by construction, checks only that semantic map, not the circuit's own layout, permutation, or phase records: a defect there, or in a fixed-seed run's determinism, can pass unseen though the record is public. This empirical software engineering study of quantum transpiler correctness uses repository mining to measure how often this happens in real merged compiler fixes: a systematically identified corpus of Qiskit transpiler bug-fixes, classified by an independently dual-coded, source-validated fault-manifestation taxonomy. Nineteen of 68 fixes (28%, 95% Wilson CI 19-40%) repair faults invisible to this equivalence screen, even one augmented with compilation-validity, circuit-quality, and performance checks, and an extended 104-fix corpus over a wider window holds at the same rate with a tighter interval (29/104, 27.9%, CI 20-37%). A conservative floor remains even restricted to the one unconditionally equivalence-invisible channel (a corrupted layout or permutation record): 10 of 68 fixes (15%, CI 8-25%) beneath the 28% headline. The gap is not Qiskit-specific: it replicates in tket (7/21, 33%), with Cirq smaller but consistent. We detected no systematic differences on five inexpensive PR-level characteristics (19 vs 49, underpowered alone; the extended 29-vs-75 corpus tightens every interval toward zero). This class dominates the invisible set, concentrating at representation-boundary crossings. We release the corpus, codebook, and coding artifacts for quantum software testing research. Here we only measure it.

Comments22 pages, 4 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑