发表机构
IBM Research; Univ. of Granada, Spain; IFMIF DONES, Spain; Instituto de Ciencia de Materiales de Madrid (ICMM); CSIC, Spain; Universidad de Alicante, Spain(IBM研究院; 格拉纳达大学; IFMIF DONES; 马德里材料科学研究所; 西班牙国家研究委员会; 阿利坎特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出eolas模块化流水线,利用大语言模型自动将科学文献转换为符合指定本体的知识图谱,并在辐照材料领域验证其高效性,同时引入首个基准数据集,通过168个实验总结出实用指南。
AI 中文摘要
对新材料的探索越来越依赖于预测模型和涵盖从原子到宏观尺度的全面模拟。然而,这些模型和模拟所需的关键数据往往以非结构化文本的形式嵌入科学文献中,限制了数据的可重用性,并给寻求有效利用现有知识的研究人员带来了挑战。虽然使用大语言模型从非结构化文本中提取结构化数据正日益流行,但传统方法通常生成具有简单模式(schema)的键值对数据。相比之下,我们引入了eolas,一个模块化流水线,它使用大语言模型自动将科学文档转换为与指定本体(ontology)对齐的知识图谱。我们展示了eolas在提取对研究旨在承受聚变反应堆中极端温度和辐射水平的材料科学家有用的信息方面的有效性。虽然人类专家可能花费三十到九十分钟从一篇文章中提取相关数据,但eolas只需几分钟就能生成高质量的知识图谱。这些图谱以表格形式呈现,并带有分面导航,便于人工验证。此外,我们引入了第一个基准数据集,旨在评估大语言模型在辐照材料领域构建知识图谱的能力。使用我们的数据集、各种大语言模型和提示技术对168个实验进行的分析提供了关键见解,我们将其总结为有效提取与输入本体对齐的知识图谱的实用指南。
英文摘要
The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels. However, essential data necessary for these models and simulations are often embedded in scientific literature as unstructured text, limiting reusability and posing challenges for researchers seeking to leverage existing knowledge effectively. While extracting structured data from unstructured text using large language models is gaining popularity, traditional methods typically generate key-value pairs data with straightforward schemas. In contrast, we introduce eolas, a modular pipeline that uses large language models to automatically transform scientific documents into knowledge graphs aligned with a specified ontology. We demonstrate eolas effectiveness in extracting useful information for scientists studying materials designed to endure the extreme temperatures and radiation levels found in fusion reactors. While a human expert might spend between thirty to ninety minutes extracting relevant data from an article, eolas can generate high-quality knowledge graphs in just a few minutes. These are presented in a tabular format with faceted navigation for easy human validation. Additionally, we introduce the first benchmark dataset designed to assess large language models capabilities in constructing knowledge graphs within the domain of irradiated materials. The analysis of 168 experiments using our dataset, various large language models and prompting techniques provides key insights that we summarize into practical guidelines for effectively extracting knowledge graphs aligned with an input ontology.