arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

APEXA:多智能体LLM自动化同步辐射数据缩减的执行完整性强制

APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction

Pawan K. Tripathi, Hemant Sharma, Andrew Chuang, Mathew J. Cherukara

arXiv 2609.24165首次发表:更新:

发表机构

Argonne National Laboratory(阿贡国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

APEXA通过确定性工具层守卫强制执行完整性,防止LLM智能体伪造未执行结果,并发布APEXA-Bench基准,在真实束线数据上验证自动化校准与积分。

AI 中文摘要

同步辐射数据缩减,即探测器校准后对太字节级衍射序列进行方位角积分,是一个多步骤、依赖专家的瓶颈,日益限制用户设施的科学产出率。LLM智能体有望消除这一瓶颈,但使用随机模型驱动真实流水线会产生聊天基准无法察觉的失败模式:智能体可能报告从未计算过的校准结果。此处的正确性是已执行操作的性质,而非对话记录的性质。我们提出APEXA,一个已部署的多智能体框架(在异构计算上运行61个工具,作为单一推理循环),在主要光源设施中通过自然语言自动化校准和积分。我们做出三项贡献。第一,执行完整性强制:一个确定性的工具层守卫,拒绝呈现任何未经执行工具调用支持的结果,其解析器能容忍跨模型的工具调用格式漂移;在部署中,一个前沿模型为从未运行的命令伪造了完整的校准比较报告,守卫将其转换为明确的非结果;同一代码门控可选电机控制表面,在模拟IOC上实现0/200次对抗性违规,而等效安全提示为15/200。第二,我们发布APEXA-Bench,一个包含58个设施任务(50个基础任务加8个跨探测器子集)的评估框架,按四类物理后果分类法组织,这是我们已知的第一个将浪费计算周期与损坏仪器区分开的基准轴;其针对NIST可溯源晶格常数的跨探测器评分揭示了两个潜在流水线缺陷。大规模智能体评分留待完整研究。第三,我们在真实束线数据上验证APEXA:从一个自然语言提示中恢复探测器几何并积分完整的衰减/曝光扫描。我们发布框架、评估框架和轨迹。

英文摘要

LLM agents can make expert scientific workflows accessible through natural language, but a plausible response does not establish that the requested computation ran. This matters at synchrotron facilities, where calibration and integration are multi-step operations over terabyte-scale datasets. We present APEXA, a deployed framework coordinating 61 tools for synchrotron data reduction. A deterministic execution-integrity guard prevents it from returning results for tool calls that did not execute, and the same tool layer enforces motor safety: across 50 adversarial scenarios and four models, it produced 0/200 violations against a simulated EPICS controller, versus 15/200 for a prompt-only baseline. We also introduce APEXA-Bench, 58 facility tasks organized by scientific correctness and physical consequence. On real APS beamline data, one prompt recovered detector geometry and integrated a full attenuation and exposure sweep. Trustworthy scientific agents need deterministic validation of execution, not model-side assurances. We release the framework, benchmark, and traces.

Comments5 pages, 4 figures. Accepted at the 4th TPC Workshop @ SC'26 (Building Open AI Infrastructure, Models, and Agentic Systems for Science). Code and benchmark: https://github.com/AdvancedPhotonSource/APEXA-APS-Beamline-Assistant ; data: https://doi.org/10.18126/tgg4-1m26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑