发表机构
Southern Oregon University; University of California, Los Angeles(南俄勒冈大学; 加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过天文学案例研究,评估开放权重离线RAG-LLM平台AquiLLM在检索和科学分析任务中的忠实性,发现其在明确检索问题上可靠,在综合与歧义消解任务上忠实性下降,凸显领域专家评估的重要性。
AI 中文摘要
科学研究日益依赖于大规模、异构的数据源,这激发了人们对检索增强生成(RAG)系统的兴趣,这些系统为科学知识和研究工作流程提供自然语言访问。研究人员正在探索这些系统作为文档搜索以及生成分析代码和流水线组件的自然语言界面的可行性。与此同时,对数据隐私和研究基础设施控制权的担忧,促使人们对开放权重模型和托管在研究机构内的开源部署产生兴趣。在天文学领域,这一发展沿袭了计算基础设施开发的悠久历史,从档案数据库和基于SQL的系统到LLM辅助研究工具。本文对AquiLLM的忠实性进行了领域专家评估,AquiLLM是一个开放权重、离线的RAG-LLM平台,旨在支持科研团队使用和保存隐性及正式知识。我们将忠实性定义为生成响应在多大程度上基于检索到的科学上下文,且不包含无根据的陈述或遗漏。我们报告了一项天文学案例研究的结果,该研究在检索和科学分析任务中评估了AquiLLM。AquiLLM在基于RAG集合的明确检索导向问题上表现最为可靠,而对于需要综合或歧义消解的查询,其忠实性会下降。这些结果凸显了开放权重RAG-LLM系统在科学研究中的前景和局限性,并证明了超越标准基准排行榜的领域专家评估的重要性。
英文摘要
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific knowledge and research workflows. Researchers are exploring the viability of these systems as natural language interfaces for document search and for generating analysis code and pipeline components. At the same time, concerns about data privacy and control over research infrastructure have motivated interest in open-weight models and open-source deployments hosted within research institutions. In astronomy, this development follows a long history of computational infrastructure development, from archival databases and Structured Query Language (SQL)-based systems to large language model (LLM)-assisted research tools. This paper presents a domain-expert evaluation of faithfulness for AquiLLM, an open-weight, offline RAG-LLM platform designed to support scientific research groups in the use and preservation of tacit and formal knowledge. We define faithfulness as the extent to which generated responses remain grounded in retrieved scientific context without unsupported claims or omissions. We report results from an astronomy case study evaluating AquiLLM across retrieval and scientific analysis tasks. AquiLLM performs most reliably on explicit retrieval-oriented questions grounded in the RAG collection, while faithfulness degrades for queries requiring synthesis or ambiguity resolution. These results highlight both the promise and limitations of open-weight RAG-LLM systems for scientific research and demonstrate the importance of domain-expert evaluation beyond standard benchmark leaderboards.
Comments14 pages, 1 figure