arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于自主量子传感实验中科学推理的智能体人工智能

Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments

Takuya Isogawa, Ryotaro Okabe, Nutdech Phadetsuwannukun, Mingda Li, Paola Cappellaro

arXiv 2607.25145首次发表:更新:

发表机构

Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究围绕大语言模型智能体实现用于金刚石中NV中心自主实验的智能体人工智能工作流程,主要贡献为展示自主实验流程,引入离线基准评估智能体推理,结果表明自主实验中智能体与代码分工明确,智能体负责假设和数据评估,代码控制硬件与执行安全约束。

AI 中文摘要

我们实现了一个围绕大语言模型(LLM)智能体构建的智能体人工智能工作流程,用于金刚石中氮空位(NV)中心的自主实验。NV中心是广泛用于量子传感的平台,计算机控制多测量的能力使NV实验成为自主工作流程的自然场景。我们有两个主要贡献。首先,展示了一个自主NV实验工作流程,结合持久项目记录、定量计算和数据分析工具以及确定性实验控制。在一次自主实验中,智能体选择单个NV中心,校准其共振频率,用拉姆齐测量法测量\(T_2^\ast\),并添加CPMG测量以检查可能与附近\(^{13}\mathrm{C}\)相关的微弱特征。其次,引入两个离线基准来分别评估智能体与实验室执行无关的推理。我们用GPT - 5.4、GPT - 5.5和GPT - 5.6 Sol评估了这两个基准。在拉姆齐检查点基准中,更多推理工作通常能更好地识别残余共振校准偏移。相比之下,在脉冲光探测磁共振(pODMR)数据评估基准中,仅脉冲序列信息在更高推理工作下会产生更多误正共振判断。要求进行预期信号计算能使所有三个模型和推理设置下的误正率保持较低。结果表明自主实验有明确分工。智能体形成科学假设并使用定量工具评估数据,而确定性代码控制硬件并执行安全约束。

英文摘要

We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond. NV centers are a widely used platform for quantum sensing, and the ability to control many measurements from a computer makes NV experiments a natural setting for autonomous workflows. We make two main contributions. First, we demonstrate an autonomous NV experiment workflow that combines persistent project records, quantitative calculation and data analysis tools, and deterministic experiment control. In one autonomous experiment, the agent selected a single NV center, calibrated its resonant frequency, measured \(T_2^\ast\) with Ramsey measurements, and added a Carr--Purcell--Meiboom--Gill (CPMG) measurement to check a weak feature that could be related to nearby \(^{13}\mathrm{C}\). Second, we introduce two offline benchmarks that evaluate the agent's reasoning separately from laboratory execution. We evaluated both benchmarks with GPT-5.4, GPT-5.5, and GPT-5.6 Sol. In the Ramsey checkpoint benchmark, greater reasoning effort generally improved recognition of a residual resonance calibration offset. By contrast, in the pulsed optically detected magnetic resonance (pODMR) data evaluation benchmark, pulse sequence information alone produced more false positive resonance judgments at higher reasoning effort. Requiring an expected signal calculation kept false positive rates low across all three models and reasoning settings. The results suggest a clear division of labor for autonomous experiments. The agent forms scientific hypotheses and uses quantitative tools to evaluate data, while deterministic code controls the hardware and enforces safety constraints.

Comments11 pages, 4 figures + Supplementary Information

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑