面向中学和高中科学主题的大语言模型理解能力基准
A Benchmark for LLM's Understanding of Middle School and High School Science Topics
浏览论文内容
中文总结 AI 辅助
针对中学科学教育缺乏标准对齐评估工具的问题,构建NGSS对齐基准并评估九个开放权重LLM,发现模型大小非性能关键,需人机协同优化项目生成。
中文摘要 AI 辅助
大语言模型(LLMs)正日益融入教育场景,然而教育工作者缺乏稳健的、与标准对齐的工具来评估其在K-12科学情境中的有效性。现有基准主要评估通用语言能力或高级科学推理,在理解LLMs对中学科学课程直接相关内容的表现方面存在关键空白。为填补这一空白,我们开发了一个全面的、对齐NGSS(下一代科学标准)的基准,覆盖初中和高中科学,采用严格的合成数据流水线、多评判者验证和项目级心理测量分析。使用该基准对九个开放权重LLM进行了系统评估,结果表明,多个较小的、可本地部署的模型在不同科学领域和问题类型上均达到了较高准确率。我们的发现表明,模型大小并非总能预测性能,这凸显了在教育部署中进行有意模型选择的重要性。随后,我们将人类评审者纳入循环,审查由LLM生成的项目与NGSS标准的一致性。人工评审显示,合成生成的项目与NGSS标准并非完全对齐,这表明了人机协同项目开发的优势、探索内容与教学知识交叉点的必要性,以及扩展基准以评估LLMs在教育场景中提供交互式、基于证据反馈的能力的需求。
英文摘要
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs' performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs' capacity for interactive, evidence-based feedback in educational scenarios.
发表机构
- University of Florida(佛罗里达大学)
- College of Staten Island, City University of New York(纽约市立大学斯塔滕岛学院)
机构由 AI 辅助整理,请以论文原文为准。