arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Alice:一个基于评分标准的多维度自动简答题评分的大规模德语基准

Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

Zhifan Sun, Sebastian Gombert, Jannik Lossjew, Tobias Wyrwich, Berrit Katharina Czinczel, David Bednorz, Marcus Kubsch, Knut Neumann, Hendrik Drachsler

arXiv 2610.09661首次发表:更新:

发表机构

DIPF | Leibniz Institute for Research and Information in Education; IPN | Leibniz Institute for Science and Mathematics Education; Umeå University; Goethe University Frankfurt(DIPF | 莱布尼茨教育与信息研究所; IPN | 莱布尼茨科学与数学教育研究所; 于默奥大学; 法兰克福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Alice,一个大规模德语自动简答题评分基准,涵盖学习表现、知识要素和技能三个子任务,并验证了评分标准检索方法及多种模型性能,发现LLM在零样本下对知识要素和技能评分困难。

AI 中文摘要

自动简答题评分(ASAS)是教育领域自然语言处理的核心。然而,公开可用的基准仍然稀缺,现有数据集大多关注学生直接回答问题的好坏,而非他们对潜在概念(如热能)或认知活动(如推理或主张)的掌握程度。为解决这一空白,我们引入了Alice,一个大规模、基于评分标准、教学法对齐的德语ASAS数据集,包含三个子任务:(i)学习表现(Alice-LP),(ii)知识要素(Alice-KE),以及(iii)技能(Alice-SK)。我们进一步将基于评分标准的ASAS形式化为评分标准检索任务,并使用一系列语言模型(从仅编码器模型到轻量级LLM)对数据集进行基准测试。我们还通过LLM的零样本提示和标准分类基线对数据集进行基准测试。实验表明,LLM在零样本设置下尤其难以对知识要素和技能进行评分。实验还表明,评分标准文本通常是有用的,尤其是对于Alice-KE和Alice-SK,而在Alice-LP上,相对于基于样本解答的输入,增益更为有限,且因模型和输入格式而异。

英文摘要

Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder-only models to lightweight LLMs. We also benchmark the dataset with zero-shot prompting via LLMs and a standard classification baseline. The experiments show that LLMs, in particular, struggle to score knowledge elements and skills in the zero-shot setting. They also indicate that rubric text is often useful, especially for Alice-KE and Alice-SK, while on Alice-LP gains over sample-solution-focused inputs are more modest and vary by model and input format.

CommentsEMNLP2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑