arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于人工智能的评分系统低估了语言薄弱学生在物理解释中的概念理解能力

AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics

Markus S. Feser, Paul L. Tschisgale

arXiv 2607.28210首次发表:更新:

发表机构

Leibniz Institute for Science and Mathematics Education(莱布尼茨科学与数学教育研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现,所有AI评分方法均存在低估语言薄弱学生物理解释中概念理解的语言偏见,类似物理教师的相关偏见,多语言学习者受影响最大。

AI 中文摘要

解释物理现象是物理学习的核心,因为学生的解释能为其概念理解提供证据。由于概念理解只能通过语言推断而非直接观察,区分概念理解与语言质量是评估的根本挑战。本研究探究基于人工智能的评分方法能否独立于学生物理解释的语言质量来评估其概念理解,将9种机器学习评分方法和2种大语言模型评分方法生成的概念理解分数,与116名中学生解释的专家人工评分进行对比。尽管与专家评分整体一致性较好,但语言质量较低的解释更易被低估,即获得的人工智能生成概念理解分数低于专家评分,这种语言偏见在所有人工智能评分方法中均存在,而语言质量较高的解释未出现类似高估情况。值得注意的是,该语言偏见与此前物理教师中报告的情况高度相似,表明难点不在于特定评估者,而在于从文本解释推断概念理解的本质。这对多语言学习者影响最大,其语言能力可能被误读为理解较弱,且随着人工智能评分用于高风险决策,影响会进一步扩大。

英文摘要

Students' explanations of scientific phenomena provide important evidence of their conceptual understanding. However, because conceptual understanding can only be inferred through language rather than observed directly, distinguishing conceptual understanding from linguistic quality represents a fundamental challenge for assessment. This challenge can also be expected to arise when artificial intelligence (AI) is used to score students' text-based explanations. We investigated whether AI-based scoring approaches assess students' conceptual understanding independently of the linguistic quality of their explanations about a phenomenon taken from a physics context. Conceptual understanding scores assigned by nine machine learning-based scoring approaches and two large language model-based scoring approaches were compared with expert-assigned scores for 116 explanations produced by secondary-school students in Germany. Multinomial logistic regression analyses examined whether expert-rated linguistic quality was associated with under- or overestimation of conceptual understanding. Across all eleven AI-based scoring approaches, lower linguistic quality was associated with a greater likelihood of underestimation, although the strength and statistical significance of this association varied between approaches. AI-based scoring approaches may thus systematically underestimate the conceptual understanding expressed in linguistically weaker explanations. This pattern resembles language bias previously documented in STEM teachers' assessment practices. The findings highlight the importance of examining not only agreement with expert ratings but also whether AI-based scoring approaches inadvertently rely on construct-irrelevant information. Addressing language bias is therefore essential for the validity and fairness of AI-supported assessment in physics education and STEM education more broadly.

CommentsShared lead authorship

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑