发表机构
East China University of Science and Technology; Shanghai Jiao Tong University; Tencent(华东理工大学; 上海交通大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对科学语言模型文献访问碎片化问题,提出证据支撑的机制知识基底MS$^3$,组织文献为证据单元与机制路径,实验证明其提升科学正确性与答案完整性。
AI 中文摘要
科学语言模型通常通过无类型的文本块访问文献,这割裂了机制丰富型问题所需的功能与证据结构。我们引入一种基于证据的机制知识基底,将科学文献组织为带来源链接的证据单元、角色类型化实体和有向机制路径。我们将其实例化为MS$^3$——一种面向导电纤维柔性传感器的材料-传感器-信号-系统模式,覆盖13,689篇论文、131,083个证据项和26,648个机制对象。在域内和覆盖偏移的问答基准上,我们比较了十种语言模型上的闭卷生成、网络搜索、原始PDF检索增强生成和MS$^3$检索。MS$^3$提升了宏平均科学正确性,同时改善了引文蕴含和答案完整性。这些结果支持机制基底作为科学语言模型的可靠表示层,并促使一种源修复工作流:当MS$^3$证据不足时,触发对其关联论文的定向检索,而非假设用户已提供正确的PDF。
英文摘要
Scientific language models often access literature through untyped text chunks, which fragment the functional and evidential structure required for mechanism-rich questions. We introduce an evidence-grounded mechanism knowledge substrate that organizes scientific literature into provenance-linked evidence units, role-typed entities, and directed mechanism paths. We instantiate it as MS$^3$, a Material-Sensor-Signal-System schema for conductive-fiber flexible sensors, over 13,689 papers, 131,083 evidence items, and 26,648 mechanism objects. On in-domain and coverage-shift question-answering benchmarks, we compare closed-book generation, Web search, Raw-PDF RAG, and MS$^3$ retrieval across ten language models. MS$^3$ improves macro-averaged scientific correctness. It also improves citation entailment and answer completeness. These results support mechanism substrates as a reliable representation layer for scientific language models and motivate a source-repair workflow in which insufficient MS$^3$ evidence triggers targeted retrieval from its linked papers rather than assuming that a user has already supplied the correct PDFs.
Comments8 pages, 5 figures, 2 tables