arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于教育评估中可扩展自动作文评分的成本效益高的生成式人工智能摘要

Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment

Haowei Hua

arXiv 2607.15829首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对自动作文评分受输入长度限制问题,提出用生成式人工智能辅助摘要框架,用GPT-5变体生成摘要作输入,结合手工语言特征形成混合框架,评估其在评分等方面表现,揭示相关权衡,为AES提供初步评估及基线等。

AI 中文摘要

自动作文评分(AES)可实现可扩展的评估和及时反馈,但仍受变压器输入长度限制的挑战,处理长作文时会导致信息丢失。本研究提出了一个生成式人工智能辅助的摘要框架,以在保持评分可靠性的同时改善长篇作文的表示。使用ASAP 2.0数据集,我们用三个GPT-5变体(GPT-5、GPT-5 mini和GPT-5 nano)生成固定长度的摘要,并将其用作下游AES模型的输入。为保留原始写作信号,从全文中提取的手工语言特征与摘要表示相结合,形成一个混合框架。该方法在评分性能、摘要质量和计算成本方面进行了评估。评分可靠性用二次加权kappa(QWK)衡量,摘要质量通过词汇重叠、语义相似性、信息保留和冗余指标进行评估。结果表明,GPT-5 mini与人类评分的一致性最高,而GPT-5的摘要质量最强。高分作文的摘要质量下降,这表明更复杂的写作在不损失信息的情况下更难压缩。这些发现揭示了模型能力、摘要保真度、成本效率和教育结构保留之间的权衡。本研究为基于GPT的AES摘要提供了初步的对照评估,并确定了未来推广所需的重要基线和消融研究。总体而言,生成式人工智能摘要为可扩展的写作评估提供了一种有前途的方法,同时需要仔细验证信息保留和公平性。

英文摘要

Automated essay scoring (AES) enables scalable assessment and timely feedback but remains challenged by transformer input-length limitations, which can cause information loss when processing long essays. This study proposes a generative AI-assisted summarization framework to improve long-form essay representation while maintaining scoring reliability. Using the ASAP 2.0 dataset, we generate controlled-length summaries with three GPT-5 variants (GPT-5, GPT-5 mini, and GPT-5 nano) and use them as inputs for downstream AES models. To preserve original writing signals, handcrafted linguistic features extracted from full essays are integrated with summary representations to form a hybrid framework. The approach is evaluated in terms of scoring performance, summarization quality, and computational cost. Scoring reliability is measured using quadratic weighted kappa (QWK), while summary quality is assessed through lexical overlap, semantic similarity, information retention, and redundancy metrics. Results show that GPT-5 mini achieves the highest agreement with human ratings, whereas GPT-5 produces the strongest summarization quality. Summary quality decreases for higher-scoring essays, indicating that more complex writing is more difficult to compress without information loss. These findings reveal trade-offs among model capacity, summary fidelity, cost efficiency, and preservation of educational constructs. This study provides an initial controlled evaluation of GPT-based summarization for AES and identifies important baselines and ablation studies required for future generalization. Overall, generative AI summarization offers a promising approach for scalable writing assessment while requiring careful validation of information preservation and fairness.

Comments23 pages, 7 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑