大型知识模型:从论文到科学推理景观
Large Knowledge Model: A Knowledge Foundation for Agentic Science at Scale
- DP Technology
- Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所)
- International Center for Quantum Materials, School of Physics, Peking University(北京大学物理学院量子材料科学中心)
- Lanzhou Center for Theoretical Physics, Lanzhou University(兰州大学理论物理研究中心)
- AI for Science Institute(AI for Science 研究院)
- School of Mathematical Sciences, Peking University(北京大学数学科学学院)
- Center for Machine Learning Research, Peking University(北京大学机器学习研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
大型知识模型将论文转化为推理图,构建科学推理景观,支持推理感知检索和问答,在三个基准上分别提升准确率9.30%、4.20%和14.69%。
AI中文摘要:
积累的科学知识在先前发现帮助研究者选择新问题、设计调查和解释结果时推动探究进展。大规模实现这一价值需要获取连接研究问题、科学程序、结论和证据的推理过程。我们引入大型知识模型(LKM),一种将文献转化为共享的、计算可访问的推理资源的科学知识基础设施。LKM将论文表示为基于来源的推理图,将结构遍历与同一对象上的语义检索耦合,并跨论文对齐相关问题、主张和推理链。这种表示形成具有三个连接视图的科学推理景观:组织研究问题和开放方向的问题景观,展示可重用科学程序的工作流景观,以及将结论与其支持、分歧和条件相连接的证据景观。统一底层支持推理感知的科学检索、基于证据的问答、比较证据分析和研究规划。研究者和智能体可以通过科学意图检索相关工作,用可检查的支持论证综合答案,并基于既有工作流和未解决证据制定研究计划。我们描述了一个语料库规模的系统,并评估了科学检索和知识密集型问答。在回答模型固定的情况下,LKM检索在ChemBench、PubMedQA和SciBench上分别将准确率提高了9.30%、4.20%和14.69%。通过将知识访问与科学推理和行动相连接,LKM为发现相关研究、重用科学知识以及协调跨研究者、智能体和研究周期的累积探究提供了共同基础。
英文摘要:
Agentic science envisions many autonomous agents investigating concurrently while building on a shared, evolving body of scientific knowledge. This requires a knowledge foundation that supports high-concurrency access, preserves traceable and reusable reasoning, and grows incrementally. We propose the Large Knowledge Model (LKM), a growing, agent-native knowledge foundation that provides a general representation of scientific knowledge across disciplines. LKM organizes the scientific literature into reasoning graphs, with claims as the core nodes and associated reasoning chains that make explicit how premises and evidence support conclusions. These source-grounded objects are persistent and addressable; cross-paper links organize them into aligned question, workflow, and evidence views. Newly extracted papers extend the foundation incrementally while preserving existing object identities. Building on this foundation, we develop an agent-native, reasoning-aware scientific retrieval system that retrieves claims together with their reasoning chains and sources, enabling agents to inspect and reuse the evidence underlying scientific conclusions. Across benchmarks, agents using LKM retrieve more evidence, cite more faithfully, and answer scientific questions more accurately: LKM nearly doubles the known supporting and contradicting evidence retrieved on SciFact-Open (818 versus 443 claim-paper pairs), reasoning graphs raise citation F1 on ScholarQABench by more than 5 points over the same retrieved papers, and LKM retrieval improves a fixed answering model by 9.3, 4.2, and 14.7 points over no retrieval on ChemBench, PubMedQA, and SciBench. LKM lays the foundation for a scientific ecosystem in which AI scientists not only recall accumulated knowledge but also extend it, returning new questions, workflows, and evidence to a memory that every subsequent investigation can build on.