发表机构
University of Georgia(佐治亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过整合生物数据库构建知识图谱,证明完整机制路径能增强LLM假设生成的证据基础,而仅端点信息虽得分高但证据弱。
AI 中文摘要
大型语言模型(LLMs)能够生成生物医学假设,但目前尚不清楚它们是基于科学证据进行推理,还是仅仅产生听起来令人信服的想法。为研究这一问题,我们将三个主要生物学数据库——京都基因与基因组百科全书(KEGG)、Rhea和UniProt——整合为一个统一的生物化学知识图谱,并构建了一个包含550条连接酶源到罕见疾病端点的路径的基准,从六个LLMs在四种条件下生成13,200个假设,每种条件改变模型接收的生物学信息:仅源酶、完整生物学路径、或仅源和疾病端点。假设使用专家推导的五标准评分细则,每个标准按1-5分制评分。我们发现,同时获得源和疾病端点的模型通常产生得分最高的假设,表明LLMs能从最少信息生成令人信服的想法。然而,这些假设在证据基础上较弱。相反,获得完整生物学路径的模型生成的假设与已知机制关系更一致。我们称之为循证约束推理。为确认这一效应,我们打乱中间路径步骤同时保持端点固定。证据基础显著下降(delta = -0.793, p < 0.001),确认模型在推理过程中确实使用了路径结构。我们的发现表明,知识图谱以两种方式支持假设生成:它们识别文献中缺失的生物学端点对,其机制路径指导LLMs如何在它们之间推理。
英文摘要
Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major biological databases: the Kyoto Encyclopedia of Genes and Genomes (KEGG), Rhea, and UniProt, into a unified biochemical knowledge graph and construct a benchmark of 550 paths connecting enzyme sources to rare disease endpoints, yielding 13,200 hypotheses from six LLMs under four conditions varying the biological information each model receives: source enzyme only, full biological path, or source and disease endpoint only. Hypotheses are scored using an expert-derived five-criterion rubric on a 1-5 scale per criterion. We find that models given both the source and disease endpoint often produce the highest-scoring hypotheses, showing that LLMs can generate compelling ideas from minimal information. However, these hypotheses are less grounded in the evidence. In contrast, models given the full biological path generate hypotheses more consistent with known mechanistic relationships. We call this evidence-disciplined reasoning. To confirm this effect, we shuffled intermediate path steps while keeping endpoints fixed. Evidence grounding dropped significantly (delta = -0.793, p < 0.001), confirming models genuinely used path structure during reasoning. Our findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.
Journal refProceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)