发表机构
Teachers College, Columbia University; New York University(哥伦比亚大学教师学院; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过JorGPT数据集分析LLM在高等教育中的评分与反馈,发现其存在子维度冗余、误解检测率低及语气缺乏调节等问题,并指出程序性主题评分更可靠,为人工监督提供了依据。
AI 中文摘要
随着大型语言模型(LLMs)在高等教育中越来越多地被用于自动评分和反馈,其结构化输出,包括多维评分量表分数、详细的反馈评论和改进建议,营造了一种全面分析性评估的表象。本研究考察这些输出是否真正提供了它们所呈现的内容。使用JorGPT数据集,包含50道开放式计算机科学问题的3,041份学生回答,由人类教师和三个商业LLM进行评分,我们识别出LLM生成的评分和反馈的表观质量与实际质量之间的三个系统性差异。子维度分数高度相关(r = 0.82-0.99,VIF高达45),提供的是冗余而非独立的诊断信息。文本反馈很少检测到学生的误解(5-7%对比教师的15-31%),其功能更像覆盖清单而非诊断工具。反馈语气无论回答质量如何都保持统一积极,缺乏人类反馈中观察到的严重程度调节。此外,评分准确性因知识领域而异,程序性主题最为可靠。这些发现为哪些LLM生成的评分和反馈方面可以依赖,哪些需要持续的人工监督提供了经验依据的指导。
英文摘要
As Large Language Models (LLMs) are increasingly adopted for automated grading and feedback in higher education, their structured outputs, including multi-dimensional rubric scores, detailed feedback comments, and improvement suggestions, create an appearance of thorough analytic evaluation. This study examines whether these outputs deliver what they appear to offer. Using the JorGPT dataset of 3,041 student responses to 50 open-ended computer science questions, scored by both human instructors and three commercial LLMs, we identify three systematic discrepancies between the apparent and actual quality of LLM-generated grading and feedback. The sub-dimension scores are highly correlated (r = 0.82-0.99, VIF up to 45), providing redundant rather than independent diagnostic information. The textual feedback rarely detects student misconceptions (5-7% vs. 15-31% for teachers), functioning as a coverage checklist rather than a diagnostic instrument. The feedback tone remains uniformly positive regardless of response quality, lacking the severity modulation observed in human feedback. Additionally, grading accuracy varies significantly by knowledge domain, with procedural topics most reliable. These findings provide empirically grounded guidance on which aspects of LLM-generated grading and feedback can be relied upon and which require continued human oversight.