发表机构
AIEd Governance(AI教育治理机构)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估了多模型OCG-PRES引导的LLM简答题评分,在996个回答上验证了高信度、效度与诊断价值,支持其作为评分辅助工具而非替代人工判断。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于或提议用于教育评分,但单模型和单次运行的评估为评估用途提供的证据有限。简答题评分需要关于信度、效度、严格性、诊断价值和失败案例的证据。本研究评估了重复的多模型OCG-PRES引导的LLM评分在简答题评估中的应用。分析使用了996个SciEntsBank回答。GPT、DeepSeek和Qianwen各自在三次独立运行中对每个回答进行评分,使用五个OCG-PRES维度:概念覆盖、关系准确性、推理完整性、矛盾控制和领域相关性。评分结果与官方二元和五类别标签进行了对比评估,并与基于答案长度、Jaccard关键词重叠、TF-IDF余弦相似度和组合传统逻辑模型的非LLM基线进行了比较。所有模型的重复运行信度均较高,GPT的ICC(3,k)=.977,DeepSeek为.992,Qianwen为.981。DeepSeek在多次运行中最为稳定。GPT在官方标签一致性方面表现最强(AUC=.909),而Qianwen更为严格,在固定阈值=3.0规则下精确率更高但召回率较低。OCG-PRES评分在五个官方类别中遵循预期的诊断模式,并在AUC和F1方面优于所有非LLM基线。重复的多模型OCG-PRES评分为LLM辅助简答题评分提供了信度、效度和诊断证据。研究结果支持将其谨慎地、基于证据地用作评分支持工具,而非取代人工判断。
英文摘要
Large language models (LLMs) are increasingly used or proposed for educational scoring, but single-model and single-run evaluations provide limited evidence for assessment use. Short-answer scoring requires evidence about reliability, validity, severity, diagnostic value, and failure cases. This study evaluated repeated multi-model OCG-PRES guided LLM scoring for short-answer assessment. The analysis used 996 SciEntsBank responses. GPT, DeepSeek, and Qianwen each scored every response across three independent runs using five OCG-PRES dimensions: concept coverage, relation accuracy, reasoning completeness, contradiction control, and domain relevance. Scores were evaluated against official binary and five-category labels and compared with non-LLM baselines based on answer length, Jaccard keyword overlap, TF-IDF cosine similarity, and a combined traditional logistic model. Repeated-run reliability was high for all models, with ICC(3,k) = .977 for GPT, .992 for DeepSeek, and .981 for Qianwen. DeepSeek was the most stable across runs. GPT showed the strongest official-label alignment by AUC (.909), while Qianwen was stricter, with higher precision but lower recall under the fixed threshold = 3.0 rule. OCG-PRES scores followed expected diagnostic patterns across five official categories and outperformed all non-LLM baselines in AUC and F1. Repeated multi-model OCG-PRES scoring provides reliability, validity, and diagnostic evidence for LLM-assisted short-answer scoring. The findings support cautious, evidence-based use as a scoring support tool rather than a replacement for human judgement.