发表机构
University of Massachusetts Amherst; University of Minnesota; Brighter Research; Adelaide University(马萨诸塞大学阿默斯特分校; 明尼苏达大学; 更光明研究公司; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究跨评分标准泛化问题,采用含评分标准无关中间表示和目标作文监督的大语言模型微调框架,在增强数据集上实验,结果表明基于特征的中间结构和受控监督可提升对未见过评分标准的泛化能力。
AI 中文摘要
自动作文评分(AES)研究主要集中在跨提示泛化,即对来自未见过提示的作文进行评分,评分标准通常保持不变。然而在实践中,教育工作者可能会在评分任务中修改甚至引入新的评分标准。本文研究跨评分标准泛化,即在一组评分标准下标注的作文上训练,在之前未见过的评分标准下评估。使用带有两个组件的大语言模型微调框架,在一个增强了学生批判性思维技能的多个评分标准定义标签的AES数据集上进行实验,发现该框架能提高泛化能力,最佳微调模型性能优于GPT-5-mini提示,仅次于GPT-5。
英文摘要
Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring task, to evaluate different aspects of essays. We study cross-rubric generalization: training on essays labeled under one set of rubrics and evaluating on previously unseen rubrics, which target different aspects of the essay. We use a Large Language Model (LLM) fine-tuning framework with two components: rubric-agnostic intermediate representations, called traits, and target-essay supervision under seen rubrics during training. On an AES dataset augmented with multiple rubric-defined labels of student critical thinking skills, we find that traits improve macro F1 by 5.0% over a baseline without traits in the hardest setting, where both target rubrics and target essays are unseen during training. We further find that increasing target-essay supervision improves performance, with our best fine-tuned open-source Llama-based model outperforming GPT-5-mini prompting by 2.1% macro F1 and trailing GPT-5 by 1.9%. These results show that trait-based intermediate structure and controlled supervision improve generalization to unseen rubrics.
CommentsPublished in AI for Education Day at SIGKDD 2026