发表机构
LORIA; Université de Lorraine; CNRS(洛林信息学及其应用研究实验室; 洛林大学; 法国国家科学研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种基于LLM的自动构建和更新MLIP模型知识图谱的流程,通过信息提取和SHACL验证循环实现迭代构建,并以Matbench Discovery排行榜为例展示其可查询性。
AI 中文摘要
在材料科学领域,许多工作致力于提供概念、术语和实体的语义表示。作为对这些工作的补充,我们报告并展示了一个流程,通过该流程可以自动构建一个关于机器学习应用于材料性能预测这一快速演进领域的知识图谱,重点关注MLIP(机器学习原子间势)。这一基于LLM的流程依赖于多个步骤,从文档和文章中的信息提取,到使用SHACL约束进行验证循环以检测和纠正错误。该流程以逐个模型为基础进行,侧重于表示的一致性,从而实现迭代构建,便于添加新模型。我们通过展示从基于Matbench Discovery排行榜中列出的模型构建的知识图谱中可以查询的几个有趣方面来说明该流程。
英文摘要
Complementing the many efforts in providing semantic representations of concepts, notions, and entities in materials science, we report and illustrate a process by which we can automatically build a knowledge graph of the fast evolving field of machine learning applied to the prediction of material properties, focusing on MLIP (Machine Learning Interatomic Potential). This LLM-based process relies on multiple steps, from information extraction in documents and articles to a validation loop using SHACL constraints to detect and correct errors. It is carried out on a model-by-model basis, focusing on the consistency of representation, therefore enabling an iterative construction where the addition of new models is facilitated. We illustrate the process by showing a few interesting aspects that can be queried from a knowledge graph built from the models listed in the Matbench Discovery leaderboard.