arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能生成的反馈效果如何?对20000多篇外语作文草稿的内在和外在评估

How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts

Steven Coyne, Diana Galvan-Sosa, Ryan Spring, Machi Shimmei, Michael Zock, Keisuke Sakaguchi, Kentaro Inui

arXiv 2607.14591首次发表:更新:

发表机构

Tohoku University; RIKEN; ALTA Institute, Computer Laboratory, University of Cambridge; CNRS, LIS, Aix-Marseille University; MBZUAI(东北大学; 理化学研究所; 剑桥大学ALTA研究所,计算机实验室; 法国国家科学研究中心,艾克斯-马赛大学语言信息处理实验室; Mohamed bin Zayed大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究外语写作中人工智能生成的书面纠正性反馈效果,通过大学外语班级近2000名学生的超20000篇草稿,从教师内在评估和学生外在反馈两角度评估,发现传统专家评估与学生反馈一致性低,强调以学习者为中心评估框架的重要性。

AI 中文摘要

本研究考察外语写作环境中的反馈,聚焦书面纠正性反馈(WCF)。大语言模型能大规模提供WCF,但使其符合教学最佳实践仍是一项持续挑战。符合事实性或相关性等标准的WCF可能仍不适用于学习情境,凸显基于学习者视角进行外在评估的必要性。我们在一个有近2000名学生的大学外语班级中部署WCF系统,收集了20000多篇草稿。从两个角度评估生成的WCF:经验丰富的英语教师使用评分标准进行内在评估,以及通过学生反馈和参与度指标进行外在评估。结果显示教师专家评分与学生反馈之间的一致性较低。这些发现表明,仅靠传统专家评估可能无法从学习者角度充分捕捉WCF的可用性或帮助性,凸显了以学习者为中心的评估框架对语言教育中基于人工智能的应用的重要性。

英文摘要

This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting criteria like factuality or relevance may still be unsuitable for learning contexts, highlighting the need for extrinsic evaluation based on the learner's perspective. We deployed WCF systems in a university-level EFL class with nearly 2,000 students, collecting over 20,000 drafts. We evaluated the generated WCF from two perspectives: intrinsic evaluation by experienced English teachers using a rubric, and extrinsic evaluation via student feedback and engagement metrics. Results revealed low alignment between teacher expert ratings and student feedback. These findings suggest that traditional expert evaluation alone may not fully capture WCF's usability or helpfulness from the learner's perspective, highlighting the importance of learner-centered evaluation frameworks for AI-based applications in language education.

CommentsPre-review version of DOI https://doi.org/10.1007/978-3-032-29788-4_35, presented at AIED 2026 Late Breaking Results. Readers are encouraged to refer to the published version

Journal refAIED CCIS 3031 (2026) 247-253

DOI:10.1007/978-3-032-29788-4_35

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑