arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量子转译器的一种防泄漏、成本感知的回归测试方法

A Leakage-Safe, Cost-Aware Regression Testing Methodology for the Quantum Transpiler

Furqan Nasir, Muhammad Arif Shah, Ifitkhar Alam

arXiv 2609.35834首次发表:更新:

发表机构

City University of Science and Information Technology (CUSIT); National University of Computer and Emerging Sciences (FAST-NUCES)(城市科学与技术大学; 国立计算机与新兴大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对Qiskit量子转译器,提出防泄漏、成本感知的回归测试选择方法,通过预先注册的评估发现多样性优先选择器优于复合评分器,并显著缩小与最强基线的差距。

AI 中文摘要

诸如Qiskit之类的量子软件开发工具包(SDK)会持续修订,单个转译器(transpiler)通道的更改就可能引入软件回归,因此持续集成(CI)必须在固定预算下从庞大且成本异构的测试套件中选择并优先安排测试。我们提出了一种针对Qiskit量子转译器的防泄漏、成本感知的回归测试选择方法,并进行了预先注册的、预算约束的评估,以避免数据泄漏。该评估使用了一个成本异构的语料库(116个单元,每单元成本差异9,945倍,T_full = 69.17秒),其预算是在读取任何测试预言标签之前根据实测成本确定的。一个透明的风险评分(risk_score)选择器并未胜过简单的测试用例优先级排序基线:其平均检测-预算AUC为0.721 [95%置信区间 0.44, 0.94],而仅基于多样性(diversity-only)的基线为0.874 [0.65, 1.00](Cliff's δ = -0.64,效应量大),基于成本/历史/变更阶段的基线为0.840,随机基线为0.821。组件消融实验表明,该复合选择器的表现不如其自身的最佳信号:多样性和新颖性(而非严重性或成本)才是有效信号。第二次独立预先注册的评估(19个变异测试事件:九个原始事件加十个新的、经过验证的算子)测试了该机制所隐含的分解后的、优先考虑多样性的选择器:它缩小了与最强基线的差距,达到可忽略的效应量(0.880对0.880,δ = -0.08),同时显著优于原始复合选择器(0.880对0.768,δ = +0.57)。来自真实Qiskit CI历史的两个经过验证的前向回归事件也证实了这一点,这些事件根据预先声明的声明范围规则按事件报告,而非汇总报告。我们如实报告了这一负面结果,作为预先注册的发现,并附有其机制。本实证软件工程研究的所有代码、数据和工件均已发布,以供复现。

英文摘要

Quantum SDKs such as Qiskit are revised continually, and a single transpiler-pass change can introduce a software regression, so continuous integration (CI) must select and prioritize tests from a large, cost-heterogeneous suite under a fixed budget. We present a leakage-safe, cost-aware regression-test-selection methodology for the Qiskit quantum transpiler, with a pre-registered, budget binding evaluation that avoids data leakage. The evaluation uses a cost-heterogeneous corpus (116 units, per-unit cost spread 9,945X, T_full = 69.17 s) whose budgets were fixed from measured cost before any test-oracle label was read. A transparent risk_score selector does not beat simple test case-prioritization baselines: mean detection-vs-budget AUC is 0.721 [95% CI 0.44, 0.94] versus 0.874 [0.65, 1.00] for diversity-only (Cliff's δ = -0.64, large), with cost/history/change-stage baselines at 0.840 and random at 0.821. A component ablation shows the composite underperforms its own best signal: diversity and novelty, not severity or cost, are the effective signal. A second, independently pre-registered evaluation (19 mutation-testing events: nine original plus ten new, verified operators) tests the decomposed, diversity-first selector this mechanism implies: it closes the gap to the strongest baseline to a negligible effect size (0.880 vs 0.880, δ = -0.08) while decisively beating the original composite (0.880 vs 0.768, δ = +0.57). Two verified forward-regression events from real Qiskit CI history corroborate it, reported per event, not pooled, under a pre-declared claim-scope rule. We report the negative result as an honest, pre-registered finding with its mechanism. All code, data, and artifacts from this empirical software engineering study are released for reproducibility.

Comments25 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑