arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14605cs.CLcs.CY

研究基于大语言模型的自动作文评分中的母语偏见:对托福作文的开放权重人工智能模型进行跨提示评估

Investigating first-language bias in LLM-based automated essay scoring: A cross-prompt evaluation of an open-weight AI-model on TOEFL essays

John Maurice Gayed

AI总结:

研究基于大语言模型的自动作文评分中母语偏见,用LoRA适配的开放权重模型Gemma-3-27B-it,在托福作文语料库测试其跨提示泛化和母语评分效果,发现有评分偏移,实现了较高等级一致性,是首次此类大规模L1公平性分析。

AI中文摘要:

本研究考察了应用于自动作文评分的LoRA适配的开放权重大型语言模型(Gemma-3-27B-it)的跨提示泛化和母语(L1)评分效果。使用在“AiAWE:一个使用LoRA适配的指令微调模型的开源大语言模型自动写作评估系统”中报告的相同模型和推理配置,该模型在来自两个提示的480篇议论文上进行了微调。我们评估了在完整的托福11语料库上的评分准确性,该语料库包含来自11种母语背景的考生在八个提示下撰写的12100篇作文,这些提示在训练期间均未出现。模型原始分数(0.5 - 5.0)被映射到ETS使用的相同三个熟练程度等级(低、中、高),以便进行直接比较。模型实现了77.79%的总体等级一致性和0.702的二次加权卡帕值,相邻等级一致性为99.98%。准确性在所有八个未见过的提示中都很稳定,与训练数据主题相关的提示没有优势,表明具有强大的跨提示泛化能力。然而,模型表现出系统性的、与L1相关的评分偏移情况。在每个熟练程度等级内,来自欧洲语言背景的作文得分始终高于来自东亚语言背景的作文,这种模式并非归因于微调数据的构成。这是对用于自动作文评分的微调开放权重大型语言模型的首次大规模L1公平性分析。

英文摘要:

This study examines the cross-prompt generalization and first-language (L1) scoring effects of a LoRA-adapted open-weight large language model (Gemma-3-27B-it) applied to automated essay scoring. Using the identical model and inference configuration reported in "AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models" (Gayed, 2026), which was fine-tuned on 480 argumentative essays from two prompts, we evaluate scoring accuracy on the full TOEFL11 corpus: 12,100 essays written by test-takers from 11 first-language backgrounds across eight prompts, none of which were seen during training. The model's raw scores (0.5-5.0) are mapped to the same three proficiency bands (low, medium, high) used by ETS, enabling direct comparison. The model achieved an overall band agreement of 77.79% and a quadratic weighted kappa of 0.702, with adjacent-band agreement of 99.98%. Accuracy was stable across all eight unseen prompts, with no advantage for prompts thematically related to the training data, indicating robust cross-prompt generalization. However, the model exhibited a systematic, L1-linked scoring offset. Within every proficiency band, essays from European-language backgrounds received consistently higher scores than essays from East-Asian-language backgrounds, a pattern not attributable to the composition of the fine-tuning data. This is the first large-scale L1 fairness analysis of a fine-tuned open-weight LLM for automated essay scoring.

补充信息

↑