arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过大语言模型集成使数学知识可解释、可访问和可互操作

Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration

Jan Range, Björn Schembera, Dominik Göddeke

arXiv 2607.24512首次发表:更新:

发表机构

Stuttgart Center for Simulation Science (SC SimTech), University of Stuttgart; Institute of Applied Analysis and Numerical Simulation, University of Stuttgart(斯图加特大学斯图加特模拟科学中心 (SC SimTech); 斯图加特大学应用分析与数值模拟研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在解决数学模型文档不符合FAIR原则及访问难等问题,通过模型上下文协议(MCP)服务器将大语言模型与MathModDB集成,实现基于认知的LLM使用,提高可解释性、可访问性并简化互操作性,通过用例展示了集成优势。

AI 中文摘要

数学模型对形式化研究问题至关重要,但其文档常不符合FAIR原则。数学模型数据库(MathModDB)等知识库通过提供数学模型的精心策划、语义丰富的表示来弥补这一差距。它基于维基数据的开源基础设施Wikibase构建,利用语义网技术支持链接开放数据、协作编辑和语义丰富的元数据存储。然而,目前访问MathModDB需要复杂的网页界面操作或SPARQL及Wikibase API的专业知识,且与实际研究数据结合仍是挑战。为克服这些限制,我们提出通过模型上下文协议(MCP)服务器将大语言模型(LLMs)与MathModDB集成,该服务器公开向量索引模式检索和基于斯坦纳树的连接规划器,结合基于对话的自然语言交互与精心策划、基于认知的知识。虽基于MathModDB实例化,但该架构可应用于其他基于Wikibase的系统。我们证明此方法能实现基于认知的LLM使用,提高模型的可解释性和可访问性,简化与外部数据库和工具的互操作性。通过连续介质力学和酶动力学领域的两个涉及数学模型的用例,说明了结合LLM的可访问性与精心策划的知识库的认知安全性的好处。

英文摘要

Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge bases such as the Mathematical Model Database (MathModDB) address this gap by providing curated, semantically rich representations of mathematical models. Built on Wikibase, the same open-source infrastructure underlying Wikidata, MathModDB utilizes Semantic Web technologies to support Linked Open Data, collaborative editing, and the storage of semantically enriched metadata, making it a domain-specific knowledge graph within the broader Wikidata ecosystem. However, access to MathModDB currently requires either navigating a complex web interface or proficiency in SPARQL and Wikibase APIs, posing significant barriers for potential users. In addition, the combination of such curated knowledge bases with actual research data stored, e.g., in Dataverse repository instances, remains a challenge. To overcome these limitations, we propose integrating Large Language Models (LLMs) with MathModDB via a Model Context Protocol (MCP) server that exposes a vector-indexed schema retrieval and Steiner-tree-based join planner, combining dialogue-based natural language interaction with curated, epistemically grounded knowledge. Although instantiated on MathModDB, the architecture can be applied to other Wikibase-based systems. We demonstrate that this approach enables epistemically grounded LLM usage, improves model explainability and accessibility beyond what the standard Wikibase interface offers, and simplifies interoperability with external databases and tools, such as Dataverse data repositories. We illustrate the benefits of combining the accessibility of an LLM with the epistemic safety of a curated knowledge base through the adaptability of the MCP protocol by two use cases involving mathematical models in the fields of continuum mechanics and enzyme kinetics.

CommentsPreprint submitted to 6th Wikidata Workshop@ISWC

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑