arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WrAFT:一种用于议论文的模块化自动写作评估系统

WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

Adnan Labib, Yixuan Huang, Jiahui Wu, John Maurice Gayed, Zheng Yuan, Qiao Wang

arXiv 2607.14524首次发表:更新:

发表机构

Department of Informatics, King’s College London; Waseda University; Beijing University of Posts and Telecommunications; University of Sheffield; Hosei University(伦敦国王学院信息学系; 早稻田大学; 北京邮电大学; 谢菲尔德大学; 法政大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出WrAFT这一用于议论文的模块化自动写作评估系统,通过模块化设计及评估多种大语言模型构建,在评分上达先进水平,人机评估反馈认可度高,且开发了公开免费的交互式用户界面。

AI 中文摘要

本研究提出了WrAFT,一种写作评估和反馈工具,能为议论文提供准确可靠的分数和有效的综合反馈。WrAFT采用模块化设计,将自动写作评估任务分为评分、表层反馈和深层反馈。构建系统时,通过直接提示和监督微调方法评估了多种大语言模型,利用了一个有480篇托福独立写作官方基准分数的专有数据集。基于基准的评估表明,WrAFT在评分方面达到了先进水平,人机评估系统生成的反馈也显示出高认可度,还开发了交互式用户界面并公开免费使用。

英文摘要

This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated writing evaluation (AWE) tasks into scoring, surface-level feedback, and deep-level feedback. In building the system, various Large Language Models (LLMs) have been evaluated, including LLaMA-3.3-70B-Instruct, GPT-4o, and Claude 3.7, through both direct prompting and supervised fine-tuning approaches. A proprietary dataset of 480 TOEFL Independent Writing essays with official benchmark scores was utilized. Benchmark-based evaluation shows that WrAFT achieves state-of-the-art performance in scoring, with a quadratic weighted kappa (QWK) of 0.84 and a root mean square error (RMSE) of 0.44 against official scores on a scale of 0-5. Human evaluation of system-generated feedback also reveals high approval ratings: 96.14 percent for surface-level feedback, 93.03 percent for deep-level macro feedback, and 94.69 percent for deep-level micro feedback. An interactive user interface has been developed for the system and is publicly available and free to use.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑