arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15630cs.HCcs.AIcs.CL

评估工具对人类和大语言模型(LLMs)是否测量相同的内容?潜在结构分析

Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

  • Weizmann Institute of Science(魏茨曼科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

Alona Strugatski, Licol Zeinfeld, Giora Alexandron

AI总结:

本研究通过探索性因子分析等方法,发现高中化学和大学入学考试定量推理部分的评估,对人类与六个多模态LLMs测量的潜在结构存在系统性差异,质疑了此类评估用于推断AI能力的效度。

AI中文摘要:

大型语言模型(LLMs)的快速发展与日益广泛的部署,使得理解其能力变得愈发重要。一种常见的方法是使用最初设计用于测量人类技能与能力的评估工具(如标准化考试)来评估LLMs,并将这些工具上的表现作为证据,用以推断LLMs在评估旨在测量的相同人类技能上具备可泛化的潜在能力。然而,从效度角度来看,此类推断要求为人类建立的观测表现与潜在构念之间的关系同样适用于LLMs。具体而言,分数解读可迁移的必要条件是,对评估的反应潜在结构具有相似性。本研究在两个教育情境(高中化学和大学入学考试的定量推理部分)中检验该条件是否成立。采用案例研究设计,将人类反应数据与六个多模态LLMs生成的反应进行比较。分析方法结合探索性因子分析、因子一致性和重抽样,以评估人类学习者与LLMs之间的潜在结构相似性。在两种评估工具中,研究发现人类与LLMs的因子结构存在系统性差异,表明所分析的评估可能对人类和LLMs测量的并非相同构念。这些发现对使用教育评估来推断AI能力的评估实践的效度提出了质疑。

英文摘要:

The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.

↑