kgsteward:用于构建、复现和维护知识图谱的工具
kgsteward: a tool for building, reproducing and maintaining distributed knowledge graphs
浏览论文内容
中文总结 AI 辅助
针对生命科学知识图谱构建与维护的挑战,提出Python工具kgsteward,通过单一配置文件在RDF存储中构建、同步和验证图谱,已在多个国际项目中应用。
中文摘要 AI 辅助
生命科学领域的合作研究项目日益需要将私有的、受禁运保护的联盟数据与公共参考数据库进行整合,以获得具有统计学意义的解释。资源描述框架(RDF)非常适合这一任务:它促进了异构数据源的集成,并允许研究人员将数据及其文档作为元数据保存在同一位置,前提是知识图谱本身在项目期间保持私有。然而,由于大多数公共资源处于持续变化的状态,科学知识图谱的开发与长期维护仍然是一项具有挑战性、劳动密集型的任务。为应对这一挑战,我们提出了kgsteward,一个Python命令行工具,它通过单个版本控制的配置文件在RDF存储库中构建和维护知识图谱。kgsteward支持多种三元组存储,使本地图谱与其外部源保持同步更新,这些外部源可能已经是RDF格式,或可即时转换,并使用SPARQL 1.1 UPDATE命令即时修改进一步导入的RDF。它还可以使用SPARQL查询验证生成的图谱,这些查询同时可作为人类用户和AI代理的使用示例。kgsteward已在SIB瑞士生物信息学研究所的多个合作项目中使用,我们展示了其在两个真实世界国际研究项目中的适用性:一个项目构建了包含化学分析和相关生物活性的植物提取物库,另一个项目则调和了用于人类代谢网络重建的公共参考资源。
英文摘要
Collaborative research projects in life sciences increasingly need to integrate private, embargoed consortium data with public reference databases in order to reach statistically meaningful interpretations. The Resource Description Framework (RDF) is well suited to this task: it facilitates the integration of heterogeneous data sources, and allows researchers to keep data and their documentation as metadata in the same place, provided the knowledge graph itself remains private during the time course of the project. Nevertheless, the development and long-term maintenance of a scientific knowledge graph remains a challenging, labour-intensive endeavour owing to the state of constant flux of most public resources. To tackle this challenge, we present kgsteward, a Python command-line tool that builds and maintains knowledge graphs inside RDF stores from a single, version-controlled configuration file. kgsteward supports multiple triplestores, keeps the local graph up-to-date with its external sources possibly already in RDF, or transformed into it on the fly, and uses SPARQL 1.1 UPDATE commands to amend further imported RDF on the fly. It can also validate the resulting graph with SPARQL queries that double as usage examples for both human users and AI agents. kgsteward has already been used in several collaborative projects at the SIB Swiss Institute of Bioinformatics, and we demonstrate its applicability in two real-world international research projects: one that builds a library of plant extracts with chemical analyses and associated bio-activities, and a second that reconciles public reference resources for human metabolic-network reconstruction.