arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

始终良好与偶尔出色:人类与机器开放式反馈质量的评分标准

Consistently Good vs. Occasionally Great: A Rubric for Open-Ended Feedback Quality from Humans and Machines

Binglin Chen, Rajarshi Haldar, Max Fowler, Matthew West, Craig Zilles

arXiv 2608.21850首次发表:更新:

AI 中文总结

本文针对入门编程开放式简答题反馈制定五项标准评分标准,对比前沿大语言模型OpenAI o1与9名助教的反馈,发现LLM平均表现优于助教但存在自我偏好偏差,为教育场景部署LLM反馈提供参考。

AI 中文摘要

为学生作业提供高质量反馈对学习至关重要,但大规模提供此类反馈仍具挑战性。本文聚焦入门编程课程中开放式简答题的反馈,目标是引导学生在重新作答时取得成功,同时不透露正确答案。我们基于教育文献制定了包含五项标准的评分标准,用于评估反馈质量:(1) 认可学生答案中的正确部分;(2) (若存在)识别至少一个缺陷;(3) 提供可操作的改进指导;(4) 适当隐藏答案;(5) 使用恰当的对话语气。利用该评分标准,我们比较了前沿大语言模型(OpenAI o1)生成的反馈与9名助教针对90份学生作答给出的反馈,由3名研究人员和一个大语言模型独立对所有反馈进行评分。结果显示,尽管一名助教常产出最佳反馈,但经人类评估,该大语言模型的平均表现始终高于助教。然而,我们还发现使用大语言模型评估反馈质量时存在显著的自我偏好偏差:该大语言模型系统性地将自身输出的评分高于人类专家的评分。研究表明,这种偏差即便在跨模型评估中仍会存在,这对采用基于大语言模型的评估的研究人员提出了重要的方法论问题。我们详细描述了助教和大语言模型的表现,分析了助教反馈质量的差异来源,并探讨了在教育场景中部署大语言模型生成的反馈的意义。

英文摘要

Providing high-quality feedback on student work is essential for learning, yet delivering such feedback at scale remains challenging. In this paper, we focus on feedback for open-ended short answer questions in introductory programming, with the goal of nudging students toward success on reattempts without revealing the correct answer. We develop a five-criteria rubric grounded in educational literature for evaluating feedback quality: (1) acknowledging correct portions of the student answer, (2) identifying at least one flaw (if present), (3) providing actionable guidance for improvement, (4) maintaining appropriate concealment of the answer, and (5) using an appropriate conversational tone. Using this rubric, we compare feedback generated by a frontier LLM (OpenAI o1) to feedback from nine teaching assistants across 90 student responses, with three researchers and an LLM independently scoring all feedback. Our results show that while one TA often produced the best feedback, the LLM demonstrated consistently higher average performance than TAs, as evaluated by humans. However, we also uncover significant self-preference bias when using LLMs to evaluate feedback quality: the LLM systematically rated its own outputs higher than human experts did. This bias, which research suggests persists even in cross-model evaluation, raises important methodological concerns for researchers employing LLM-based evaluation. We provide detailed characterization of both TA and LLM performance, analyze sources of variance in TA feedback quality, and discuss implications for deploying LLM-generated feedback in educational settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑