arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02056cs.CL

HyGRAIL:面向知识图谱的成本感知且证据支撑的科学假设发现框架

HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

Yihang Sun, Zhihan Zhu, Zhiyuan Jiang, Jingyi Ge, Zixuan Li, Jiaxuan You

AI总结:

HyGRAIL是结合异构GNN分类与LLM审查的科学假设发现框架,在MatKG上F1达0.429,降LLM调用率54.36%,紧凑图证据更利于假设验证。

AI中文摘要:

科学知识图谱组织了从科学文献中提取的实体与关系,但本质上仍存在不完整性。此类图谱中缺失的类型化链接可代表合理的科学假设,例如材料与应用之间未被探索的关联。然而,科学假设发现颇具挑战性,因为在类型化候选对中,真正的发现极为稀少:图神经网络(GNN)效率高,但对模糊案例不可靠;而大型语言模型(LLM)知识丰富,但全面应用成本过高,且并非天然基于图结构。我们提出HyGRAIL,这是一种成本感知且证据支撑的框架,将异构GNN分类与基于LLM的假设审查相结合。HyGRAIL首先使用GNN对候选假设进行评分,并识别经验证校准的模糊区域,仅将图不确定的案例路由至LLM审查。对于每个被路由的假设,HyGRAIL从知识图谱(KG)中检索节点级关联与多跳关系路径,随后通过基于模板或基于LLM的自然化将此结构化证据转换为自然语言。最后,LLM审查智能体利用自然化证据与经验证选择的决策标准对每个困难假设进行判断。在MatKG数据集上,HyGRAIL实现了0.429的最佳F1分数,相较于最强的现有基准提升了0.242个F1点,相较于仅使用GNN的基准提升了0.322个F1点。与此同时,GNN分类平均降低了54.36%的LLM调用率。消融研究进一步表明,检索到的图证据对于可靠的假设验证至关重要,且紧凑的双面证据比单纯增加检索数量更有效。

英文摘要:

Scientific knowledge graphs organize entities and relations extracted from scientific literature, but they remain inherently incomplete. Missing typed links in such graphs can therefore represent plausible scientific hypotheses, such as unexplored associations between materials and applications. However, scientific hypothesis discovery is challenging because true discoveries are extremely sparse among typed candidate pairs: graph neural networks (GNNs) are efficient but unreliable for ambiguous cases, while large language models (LLMs) are knowledgeable but too costly to apply exhaustively and are not naturally grounded in graph structures. We propose HyGRAIL, a cost-aware and evidence-grounded framework that combines heterogeneous GNN triage with LLM-based hypothesis review. HyGRAIL first uses a GNN to score candidate hypotheses and identify a validation-calibrated ambiguous region, routing only graph-uncertain cases to LLM review. For each routed hypothesis, HyGRAIL retrieves node-level associations and multi-hop relational paths from the knowledge graph (KG), then converts this structured evidence into natural language through template-based or LLM-based naturalization. An LLM review agent finally judges each hard hypothesis using the naturalized evidence and validation-selected decision criteria. On MatKG, HyGRAIL achieves the best F1 score of 0.429, improving over the strongest prior baseline by 0.242 F1 points and over the GNN-only baseline by 0.322. Meanwhile, GNN triage reduces the LLM call rate by 54.36% on average. Ablation studies further show that retrieved graph evidence is crucial for reliable hypothesis verification and that compact, two-sided evidence is more effective than simply increasing retrieval quantity.

↑