发表机构
Advanced Photon Source, Argonne National Laboratory(先进光子源,阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对科学设施操作知识难覆盖问题,提出APS - RAG平台,融合多种检索通道并添加校正智能循环,构建APS - Bench数据集评估。结果显示该平台提升关键信息召回率,交叉编码器重排器作用显著,还发布相关资源,为设施操作提供可信AI辅助工作流程。
AI 中文摘要
科学用户设施积累了数十年的操作知识,没有单一搜索索引能涵盖,如电子日志、技术文档等。我们提出了APS - RAG(先进光子源检索增强生成),这是一个已部署的平台,通过自然语言查询使先进光子源的机构知识可供工作人员使用,并进行基于操作的评估。检索引擎融合密集、稀疏和知识图谱通道,采用查询类型自适应倒数排名融合,添加校正智能循环,并在模型上下文协议工具层上运行原生工具ReAct执行器。我们构建了APS - Bench,一个有50个问题的问答数据集及可审计的黄金答案。每个检索增强变体在严格关键信息召回率上比朴素的BM25基线有数值提升,完整的校正智能图谱RAG评分更高。交叉编码器重排器对答案质量有显著贡献,移除它会大幅降低严格关键信息召回率。图谱通道和校正循环有积极贡献但性能提升有限。此外,我们还比较了开源和闭源大语言模型在最终答案合成中的性能。我们发布了APS - Bench构建方法、六层评估工具和底层代码库以及‘/aps - rag’检索代理技能框架,以支持其他设施的复制和采用。该已部署平台及其基于操作的评估为设施操作中可信的、基于统计的人工智能辅助提供了一个有前景的工作流程,可转移到其他大型科学仪器上。
英文摘要
Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augmented Generation, a deployed platform that makes the institutional knowledge at the Advanced Photon Source (APS) accessible to staff through natural-language queries, along with an operations-grounded evaluation. The retrieval engine fuses dense, sparse, and knowledge-graph (KG) channels with query-type-adaptive reciprocal-rank fusion, adds a corrective agentic loop, and runs a native-tool ReAct executor over a Model Context Protocol (MCP) tooling layer. We construct APS-Bench, a 50-question, question-answering (QA) dataset with auditable gold answers. Every retrieval-augmented variant numerically improves strict vital-nugget recall over a naive BM25 baseline (63.8%), with the full corrective Agentic GraphRAG scoring (70.3%). The cross-encoder reranker contributes significantly to answer quality: removing it and allowing the LLM to score relevance drastically reduces strict vital recall by 32.8%. The graph channel and corrective loop contribute positively as expected, but the performance gains are marginal. Additionally, we also compare the performance of open-source and closed-source LLMs in final answer synthesis. We release the APS-Bench construction methodology, the six-layer evaluation harness, and the underlying codebase, along with the '/aps-rag' retrieval agent skill framework, to support reproduction and adoption at other facilities. Together, the deployed platform and its operations-grounded evaluation present a promising workflow for trustworthy, statistically grounded AI assistance in facility operations, transferable to other large scientific instruments.