发表机构
Weizmann Institute of Science(魏茨曼科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究结合探索性因子分析与学科专家盲审,发现6个LLMs的评估响应潜在因子与人类认知结构的可解释性存在显著差异,多数LLM因子无法被人类专家解读,揭示其机制与人类推理不同。
AI 中文摘要
对大语言模型(LLMs)的评估高度依赖人类设计的评估任务,这隐含假设了AI与人类采用相似的潜在认知结构。为挑战这一假设,本研究探究了决定LLM性能的潜在因素是否具备与人类学习者认知结构相同的、可被人类解释的实质性意义。我们收集了人类与6个LLMs在定量推理和化学评估中的响应,分别对两组数据进行探索性因子分析(EFA)。随后,学科主题专家(SMEs)对生成的因子图进行盲审,为浮现出的结构赋予教学意义。专家成功解读了多数人类来源的因子,但无法为定量推理任务中任何LLM来源的因子赋予意义,仅解读了化学任务中一半的LLM因子。该结合数据驱动的EFA与盲审专家解读的框架表明,LLMs常基于与人类推理不同的、统计层面上不可解释的机制运行。
英文摘要
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.
CommentsAccepted for publication at AIME 2026