arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量子工程中已验证自主性的评估

Evaluating Verified Autonomy in Quantum Engineering

Naixu Guo, Changhao Li, Siyu Cheng, Qicheng Tang, Binzhao Luo, Bikun Li, Yuxuan Du, Shihao Ru, Jiaqi Cai

arXiv 2609.17439首次发表:更新:

发表机构

Centre for Quantum Technologies, National University of Singapore; Unitary Foundation; Boston College; Georgia Institute of Technology; GaugeForge PTE. LTD.; University of Chicago; Nanyang Technological University(新加坡国立大学量子技术中心; 幺正基金会; 波士顿学院; 佐治亚理工学院; GaugeForge私人有限公司; 芝加哥大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对量子工程中智能体可靠性未验证的问题,提出Quantum-Harbor虚拟实验室和QIQCBench基准,通过49个任务评估17个系统,揭示能力展示与可靠操作间的显著差距。

AI 中文摘要

可靠的量子工程对于将量子现象转化为实用技术至关重要。随着量子平台规模和复杂性的增长,其表征和操作需要越来越多的人力投入和协调。科学人工智能智能体能够规划实验、操作仪器和分析观测结果,为自主量子工程提供了一条有前景的途径。然而,当前智能体在此环境中能否可靠执行尚未得到系统验证。为填补这一空白,我们开发了Quantum-Harbor,一个虚拟实验室,为智能体与量子系统交互提供受控执行环境。该设计能够直接验证所采取的行动和得出的结论。基于此框架,我们引入了QIQCBench,一个包含49个专家编写任务的基准,涵盖校准与控制、纠错与编译、传感与组网等多个层面。在17个前沿智能体系统中,QIQCBench揭示了已验证性能的巨大差异。这些结果暴露了展示能力与实现可靠操作之间的显著差距,并将Quantum-Harbor确立为衡量量子工程中已验证自主性进展的基础。

英文摘要

Reliable quantum engineering is essential for turning quantum phenomena into practical technologies. As quantum platforms grow in scale and complexity, their characterization and operation require increasing human effort and coordination. Scientific artificial intelligence agents, which can plan experiments, operate instruments, and analyze observations, offer a promising route towards autonomous quantum engineering. Yet whether current agents can perform reliably in this setting has not been systematically established. To fill this gap, we developed Quantum-Harbor, a virtual laboratory that provides a controlled execution environment for agents to interact with quantum systems. This design enables direct verification of both the actions taken and the conclusions drawn. Building on this framework, we introduce QIQCBench, a benchmark of $49$ expert-authored tasks spanning multiple layers including calibration and control, error correction and compilation, sensing and networking. Across $17$ frontier agentic systems, QIQCBench reveals wide variation in verified performance. These results expose a substantial gap between demonstrating capability and achieving reliable operation, and establish Quantum-Harbor as a foundation for measuring progress towards verified autonomy in quantum engineering.

Comments10 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑