arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型在数学推理中是否展现出连贯的知识结构?——来自知识空间理论的视角

Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

Peng Cui, Heejin Do, Mrinmaya Sachan

arXiv 2609.05245首次发表:更新:

发表机构

ETH Zürich; ETH AI Center(苏黎世联邦理工学院; ETH人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究基于知识空间理论构建框架,对比8款LLM与人类学习者,发现LLM无类人连贯知识结构,彼此间知识结构也不一致,且结构缺陷难以被常规评估检测。

AI 中文摘要

人类知识本质上是结构化且相互关联的:掌握一个概念需要先掌握其前提条件,这一原则由知识空间理论(KST)正式确立。尽管大语言模型(LLM)在复杂推理任务上表现出色,但它们是否展现出与人类相似的连贯知识结构仍不明确。我们引入了一个基于KST的框架,用于评估LLM在数学推理中的知识结构,并将其作为规范框架分析LLM行为是否遵循原则性的知识依赖关系。我们将8个开源和闭源LLM与真实人类学习者进行对比评估,发现:(1)LLM并不遵循人类知识结构——它们经常违反知识依赖关系,且无法利用上下文提供的相关知识来提升对依赖问题的表现;(2)LLM之间也不共享一致的知识结构,这体现为它们知识分布的重叠度较低。此外,这些结构缺陷在基于准确率和LLM作为评判者的评估中大多不可见。总体而言,我们的结果提供了行为证据,表明当前LLM的知识并不遵循人类式的结构。

英文摘要

Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of its prerequisites, a principle formalized by Knowledge Space Theory (KST). While LLMs achieve strong performance on complex reasoning tasks, it remains unclear whether they exhibit coherent, human-like knowledge structure. We introduce a KST-grounded framework for evaluating LLM knowledge structure in mathematical reasoning, using it as a normative framework to analyze whether LLM behavior adheres to principled knowledge dependencies. Evaluating eight open- and closed-source LLMs against real human learners, we find that (1) LLMs do not adhere to human knowledge structure -- they frequently violate knowledge dependencies and fail to leverage related knowledge provided in context to improve performance on dependent questions; (2) LLMs do not share a consistent knowledge structure among themselves, as reflected by low overlap in their knowledge distributions. Furthermore, these structural deficiencies remain largely invisible to accuracy-based and LLM-as-judge evaluations. Together, our results provide behavioral evidence that current LLMs knowledge does not follow a human-like structure.

CommentsEMNLP 2026 findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑