AI 中文总结
研究使用 Bugs4Q 对修复 Qiskit 程序的大语言模型进行基准测试,评估两个含 67 个缺陷的 Bugs4Q 版本,在六个 Qiskit 版本上测试四个大语言模型,发现问题并发布重新验证的基准测试,强调基准验证先于修复评估。
AI 中文摘要
在量子程序中,Bugs4Q 是一个广泛使用的包含真实量子缺陷的基准测试。然而,其评估假定基准标签保持有效且生成的修复在目标环境中执行。我们评估了两个包含 67 个独特真实 Qiskit 缺陷的 Bugs4Q 版本,补充缺失的可执行测试,并在六个固定的 Qiskit 版本上重新运行所有条目。发现量子基准测试可能存在无声标签反转问题。我们评估了四个大语言模型,GPT - 5.4 达到最高的 pass@10(48.8%)。所有模型在 Qiskit 0.45.0 上表现最佳,在 Qiskit 1.0 之后下降。许多失败源于废弃或不兼容的 API。我们发布了重新验证、版本固定的 Bugs4Q 基准测试,并表明基准验证必须先于修复评估。
英文摘要
In quantum programs, Bugs4Q is a widely used benchmark containing real quantum defects. However, its evaluation assumes that benchmark labels remain valid and that generated fixes execute in the target environment. We evaluate two Bugs4Q versions containing 67 unique real Qiskit defects, adding executable tests where missing, and re-run all entries across six pinned Qiskit releases (0.25.0, 0.45.0, 1.0.0, 1.1.1, 2.0.0, and 2.3.1). We find that quantum benchmarks can suffer from silent label inversion: entries become invalid without errors when reference fixes stop executing or buggy programs no longer reproduce failures. Thus, correctness depends on the (benchmark, version) pair rather than the benchmark alone. We evaluate four LLMs (GPT-4o-mini, GPT-5o-mini, GPT-5.4, and GPT-5.4-mini), generating up to 10 repair candidates per defect and testing them across all versions. GPT-5.4 achieves the highest pass@10 (48.8%), followed by GPT-5.4-mini (47.3%), GPT-5o-mini (30.3%), and GPT-4o-mini (22.6%). All models perform best on Qiskit 0.45.0 and decline after the Qiskit 1.0 transition. Many failures arise from deprecated or incompatible APIs rather than incorrect repairs, and 64\% of successful repairs occur on entries invalid under the target version. We release a re-validated, version-pinned Bugs4Q benchmark and show that benchmark validation must precede repair evaluation.
Comments4 pages, 1 figure