arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于大规模写作评估的AI辅助评分的人在回路框架

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

María Eugenia Curi, Germán Capdehourat, Isabel Amigo, Magdalena Romano, Rosana Serra, Adrián Silveira, Andrés Peri

arXiv 2609.05143首次发表:更新:

发表机构

Ceibal; ANEP(塞瓦尔机构; 乌拉圭全国公共教育管理局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种用于大规模写作评估的人在回路AI辅助评分框架,基于全国考试数据验证其可行性,该框架可高效分配专家精力,需结合人工监督安全应用。

AI 中文摘要

将人工智能(AI),尤其是大语言模型(LLMs)整合到教育评估中,为提升评分流程的效率和可扩展性开辟了新机遇。本研究设计并验证了一种用于大规模全国性评估中书面回答的AI辅助评分框架,该方法聚焦于约150-200词的短篇文本,采用人在回路(human-in-the-loop)策略,在保证评估质量的同时减少人工工作量。研究基于真实运营场景,使用了全国性考试近两批的数据,每批包含约5000份学生回答。我们分析了AI生成的分数与人工评分者在多个评分维度上的一致性,以及所提出的决策流程对合格/不合格结果的影响。结果显示,在大多数维度上,模型与人工评估之间存在中等到高度的一致性,支持该场景下AI辅助的可行性。此外,所提出的修正工作流能识别出最需要人工审查的案例,实现专家精力的更高效分配。研究结论表明,仅当与精心设计的人工监督结合时,AI辅助评分才能安全地整合到大规模评估流程中。论文最后讨论了在全国评估系统中部署的实际意义,并概述了未来研究方向,包括对模型-人类一致性的纵向监测,以及分析AI辅助审查工作流可能引入的认知偏差。

英文摘要

The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.

Comments25 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑