arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25607cs.AI

ArticleMiner:本体引导的科学出版物知识图谱构建

ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications

Md Abrar Jahin, Craig A. Knoblock, Jay Pujara

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ArticleMiner框架,通过人工任务模块和证据协调,从科学论文及补充文件中构建知识图谱,在四个领域任务上优于少样本基线,并引入新地球化学基准。

中文摘要 AI 辅助

科学论文的大量定量内容保存在表格和补充文件中,其中数字的含义仅通过其表头、标题、单位、分析方法以及所在领域的惯例来确定。因此,恢复表格的行和列并不等同于恢复其所报告的科学事实。大多数语义表格解释方法假设已有一个干净的表格,随后将其单元格或列映射到本体术语,而大多数出版物级提取系统则针对单一领域设计。我们研究了一条中间路径:一个共享流程,读取论文及其补充文件,从多个解析器和语言模型中收集证据,并协调这些证据,同时为每项任务提供领域含义的、有限的人工编写任务模块。该模块列出了图可能使用的规范名称、映射到这些名称的表面形式、一组简短的推导规则和有效性约束、身份键以及用于编写RDF的绑定。它定义了任务允许输出的内容;它并不试图列出某个领域的每一条惯例。我们在ArticleMiner框架中构建了四个此类模块(分别用于药物发现化学、材料科学、机器学习和矿物地球化学),并在163篇论文上进行了评估,其中包括一个新的人工专家策展的地球化学基准。在与同LLM少样本基线的比较中,点估计在全部四项任务上均有利于ArticleMiner,但在两个较小的基准上存在不确定性。地球化学比较还包括对补充文件的访问,因此其改进不能仅归因于领域指导。

英文摘要

Scientific papers keep much of their quantitative content in tables and supplementary files, where a number means something only through its header, caption, unit, analytical method, and the conventions of its field. Recovering the rows and columns of a table is therefore not the same as recovering the scientific fact it reports. Most semantic table-interpretation methods assume that a clean table is already available and subsequently map its cells or columns to ontology terms, whereas most publication-level extraction systems are designed for a single domain. We study a middle path: a shared process that reads a paper and its supplementary files, gathers evidence from several parsers and a language model, and reconciles that evidence, while a bounded human-authored task module for each task supplies the domain meaning. The module lists the canonical names the graph may use, the surface forms that map to them, a small set of derivation rules and validity constraints, an identity key, and the bindings used to write RDF. It defines what a task is allowed to emit; it does not try to list every convention of a field. We build four such modules (for drug-discovery chemistry, materials science, machine learning, and mineral geochemistry) in the ArticleMiner framework, and evaluate them on 163 papers, including a new geochemistry benchmark with expert-curated ground truth. In comparisons against a same-LLM few-shot baseline, the point estimates favor ArticleMiner on all four tasks, with uncertainty on the two smaller benchmarks. The geochemistry comparison also includes access to supplementary files, so its improvement cannot be attributed to domain guidance alone.

发表机构

  • USC Information Sciences Institute(南加州大学信息科学研究所)
  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑