arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QuSema:通过量子知识增强的智能体检测量子库中的静默缺陷

QuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced Agents

Yujin Song, Kaining Zhang, Qixin Zhang, Shuai Wang, Pingchuan Ma, Yuxuan Du

arXiv 2610.10258首次发表:更新:

发表机构

Nanyang Technological University; Hong Kong University of Science and Technology; Zhejiang University of Technology(南洋理工大学; 香港科技大学; 浙江工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

QuSema是一个利用量子语义和文档作为预言机的自主测试智能体,通过智能体循环检测量子库中的静默缺陷,在基准上优于现有工具并发现多个新缺陷。

AI 中文摘要

量子库现已成为量子算法开发的关键基础设施,然而其正确性仍然难以测试。现有的测试技术主要依赖于基于失败或基于比较的预言机,仅在执行失败、违反运行时检查或与另一实现不一致时暴露缺陷。当合适的基于执行的预言机不可用时,其适用性受到限制,导致一些静默缺陷未被检测到。此类遗漏的缺陷可能产生错误结果,并传播到实验结论、模拟研究和算法设计中。在此,我们提出QuSema,一个用于发现量子库中静默缺陷的自主测试智能体。QuSema利用量子语义和文档中的约束作为源代码级语义预言机,以评估实现逻辑是否能从有效输入产生无效输出。它通过一个智能体循环运作,反复检查库API文档和源代码,推理量子操作的预期行为,识别潜在的语义偏差,并通过库API生成可执行测试来验证它们。在量子领域推理的引导下,QuSema将高层行为不匹配转化为具体的、用户可触发的缺陷报告,使其能够发现非崩溃缺陷。我们为Qiskit和PennyLane实现了QuSema。在一个包含20个历史静默缺陷的基准上,QuSema实现了比Claude Code和Codex更高的平均缺陷定位数量,其中DeepSeek配置的成本低于Claude Code。QuSema还发现了40个先前未知且经开发者确认的缺陷,其中包括30个静默缺陷。

英文摘要

Quantum libraries are now critical infrastructure for quantum algorithm development, yet their correctness remains difficult to test. Existing testing techniques mainly rely on failure-based or comparison-based oracles, exposing bugs only when executions fail, violate runtime checks, or disagree with another implementation. Their applicability is limited when suitable execution-based oracles are unavailable, leaving some silent bugs undetected. Such missed bugs can produce incorrect results that propagate into experimental conclusions, simulation studies, and algorithmic designs. Here we present QuSema, an autonomous testing agent for finding silent bugs in quantum libraries. QuSema uses constraints from quantum semantics and documentation as a source-level semantic oracle to assess whether implementation logic can produce invalid outputs from valid inputs. It operates through an agentic loop that repeatedly inspects library API documentation and source code, reasons about the intended behavior of quantum operations, identifies potential semantic deviations, and validates them by generating executable tests through library APIs. Guided by quantum-domain reasoning, QuSema turns high-level behavioral mismatches into concrete, user-triggerable bug reports, enabling it to uncover non-crash defects. We implement QuSema for Qiskit and PennyLane. On a benchmark of 20 historical silent bugs, QuSema achieves higher mean bug relocation counts than Claude Code and Codex, with the DeepSeek configuration costing less than Claude Code. QuSema also discovers 40 previously unknown bugs confirmed by the developers, including 30 silent bugs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑