评估LLM生成的学生写作反馈中的反馈焦点与教学适应性
Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing
浏览论文内容
中文总结 AI 辅助
本研究通过构建FeedType基准,评估了六种LLM在三种提示策略下生成反馈的焦点类型与适应性,发现其虽覆盖多数焦点类型但分布不均且适应性不及专家教师。
中文摘要 AI 辅助
我们研究了最先进的大语言模型(LLMs)生成的反馈是否在反馈焦点和适应性方面反映了专家教师的教学实践。以往的评估工作考察了反馈特征、其对学习的影响及其目标,但反馈的焦点及其适应性在很大程度上仍被忽视。为弥补这一空白,我们采用并细化了Narciss的分类法,将其划分为七种反馈焦点类型,用于标注三门大学写作课程中教师和LLM生成的反馈。我们发布了FeedType基准,包含来自六种LLM在三种提示策略下生成的带标注的教师和LLM反馈。我们评估了反馈焦点类型的覆盖范围和分布,并考察LLM是否像专家教师那样根据草稿阶段和学生表现水平调整其反馈。我们的研究结果表明,尽管大多数LLM覆盖了大多数反馈焦点类型,但它们未能反映教师的反馈分布,并表现出不同程度的适应性,且没有一种LLM能与教师的适应性行为相匹配。我们认为FeedType将支持未来关于LLM反馈生成中教学对齐的研究。
英文摘要
We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined feedback characteristics, its impact on learning, and its target, yet the focus of feedback and its adaptivity remains largely overlooked. To bridge this gap, we adopt and refine Narciss's taxonomy into seven feedback focus types to annotate teacher and LLM-generated feedback across three university writing courses. We release FeedType, a benchmark containing annotated teacher and LLM feedback from six LLMs under three prompting strategies. We assess the coverage and distribution of feedback focus types, and examine whether LLMs adapt their feedback across draft stages and student performance levels as an expert instructor does. Our findings show that while most LLMs cover most feedback focus types, they fail to reflect teacher feedback distributions and show varying levels of adaptivity, with none matching the teachers' adaptive behavior. We believe FeedType will support future research on pedagogical alignment in LLM feedback generation.
发表机构
- University of Pittsburgh(匹兹堡大学)
机构由 AI 辅助整理,请以论文原文为准。