arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

错误监督的合成学习者写作用于自动作文评分

Error-Supervised Synthetic Learner Writing for Automated Essay Scoring

Duy Anh Nguyen

arXiv 2609.23573首次发表:更新:

发表机构

University of Greenwich; FPT University(格林威治大学; FPT大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种错误监督的合成作文生成方法,通过微调LLM生成器于错误标注文本,在较大数据下优于常规合成基线,并产生类似学习者的错误分布。

AI 中文摘要

合成作文有助于减少自动作文评分(AES)中对人工撰写数据的依赖。然而,它们往往缺乏真实的错误,限制了其代表真实人类写作的能力,尤其是当目标文本旨在模仿语言学习者所写的文本时。在本研究中,我们提出了一种简单的方法,将错误监督引入合成作文生成中。具体来说,我们在常用于语法错误检测(GED)的错误标注文本上微调一个大型语言模型(LLM)生成器。为了评估所提出方法的实用性,我们在三种数据条件下微调和评估AES评分器:真实作文、以常规方式生成的合成作文,以及使用我们提出的方法生成的合成作文。结果表明,在较大数据设置下,所提出的方法在12个数据集-指标比较中的11个中优于常规合成基线,其性能在某些情况下接近在真实作文上训练的模型。尽管有这些提升,在极低资源设置下的性能仍然参差不齐,相对于常规基线的优势仅在200篇训练作文时才变得更加明显,尽管并非在所有数据集上一致。定性和定量分析进一步表明,所提出的方法产生的学习者式错误的分布与真实作文中观察到的错误分布大致相似。

英文摘要

Synthetic essays can help reduce dependence on human-written data in Automated Essay Scoring (AES). However, they often lack realistic errors, limiting their ability to represent authentic human writing, particularly when the target texts are intended to resemble those produced by language learners. In this study, we present a simple approach that introduces error supervision into synthetic essay generation. Specifically, we fine-tune an LLM generator on error-annotated texts of the kind commonly used in Grammatical Error Detection (GED). To assess the utility of the proposed approach, we fine-tune and evaluate AES scorers under three data conditions: authentic essays, synthetic essays generated conventionally, and synthetic essays generated using our proposed approach. The results show that in the larger-data settings, the proposed approach outperforms the conventional synthetic baseline in 11 out of 12 dataset-metric comparisons, with performance in some cases approaching that of models trained on authentic essays. Despite these gains, performance under extremely low-resource settings remains mixed, with advantages over the conventional baseline only becoming more apparent at 200 training essays, although not consistently across datasets. Qualitative and quantitative analyses further show that the proposed approach produces learner-like errors whose distributions broadly resemble those observed in authentic essays.

Comments17 pages, 1 figure, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑