AI 中文总结
本研究探究用LLMs分析学生计算物理问题书面解释,发现其可复刻人工评估结果并规模化检测计算思维发展趋势,为大招生规模物理课程的CT评估提供可行方法。
AI 中文摘要
随着计算在物理教育中愈发核心,创建可扩展的方法来评估学生真实的计算思维(CT)仍是关键挑战。学生撰写的回答能捕捉到细致的推理,但难以规模化评估。本研究探究了大型语言模型(LLMs)在分析学生对计算物理问题的书面解释中的应用,这些解释来自学期初和学期末的调查。首先,我们基于计算思维文献建立了人工编码的基准,发现学生在数据实践和计算问题解决实践方面存在显著发展。当面对相同的回答时,LLM成功复刻了人工评估结果,并能在大规模数据集上规模化检测这些关键趋势。值得注意的是,人工评分者和LLM都难以可靠评估更复杂的构念,如系统思维。总体而言,本研究表明,LLM为大规模评估大招生规模物理课程中学生的计算思维提供了可行方法。
英文摘要
As computation becomes more central to physics education, creating scalable methods to assess authentic computational thinking (CT) in students remains a critical challenge. While student-written responses capture nuanced reasoning, they are difficult to evaluate at scale. In this study, we investigated the use of Large Language Models (LLMs) to analyze students' written explanations of computational physics problems on a pre- and post- semester survey. By first establishing a human-coded baseline, grounded in CT literature, we identified significant growth in Data Practices and Computational Problem-Solving Practices. When given the same responses, an LLM successfully mirrored the human evaluations and scaled up the detection of these key trends across a large dataset. Notably, both human raters and the LLM struggled to reliably evaluate more complex constructs such as Systems Thinking. Overall, this study demonstrates that LLMs offer a viable method to scale the evaluation of students' CT in large-enrollment physics courses
Comments1 figure, 2 tables. Submitted to Physics Education Research Conference (PERC) 2026