arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14284cs.CV

视觉语言模型在成果导向教育中手写考试标准级评分中的应用

Vision-Language Models for Criterion-Level Grading of Handwritten Examinations in Outcome-Based Education

Asif Hasan Tonmoy, Saad Ahmed, Md Khalid Syfullah, S. M. Jahangir Alam

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估视觉语言模型在手写考试标准级评分中的性能,发现Qwen2.5-VL结合LoRA优于人类一致性,并强调需针对评分标准校准、确保可重复性及审查后果性错误。

中文摘要 AI 辅助

标准级评分将考试成绩与学习成果联系起来,但人工评分带来了工作量和评分者之间的差异。本研究评估了视觉语言模型(VLMs)在手写成果导向评估中的表现,涵盖五个维度:准确性、与人类评分的一致性、重复运行的可靠性、错误集中度以及解释质量。利用来自485份本科生考试答案的1,982条标准级记录,我们比较了涵盖Qwen2.5-VL、InternVL3、Pixtral、Donut基线和级联集成的20种配置。评估设置包括零样本提示、少样本提示、部分微调和低秩适配(LoRA)。两位独立的教师评分者对全部291条测试标准进行了重新评分,提供了基于相同评估材料的人类一致性基线。Qwen2.5-VL结合LoRA在二次加权卡帕系数(QWK)上达到0.727,与考官相比的平均绝对误差为0.435分,而人类评分者对的平均QWK为0.551。这一比较反映了对考官训练分数的校准。对于所有三个指令微调的VLM,LoRA均优于部分微调,而在所有具有有效提示分数的配置中,少样本提示降低了QWK。聚合可靠性和精确可重复性出现分歧:组内相关系数范围从0.790到0.874,但在五次采样运行中,50.2%至63.6%的标准分数发生了变化。注意力引导删除相对于随机掩蔽没有显示出统计学上的显著优势,四位教师评审者对解释的有用性也未达成共识。这些发现强调了需要针对评分标准进行校准、可重复的评分、对后果性错误的审查以及对解释的单独验证。发布的评估协议支持标准级评估研究和在教师监督下的评分工具。

英文摘要

Criterion-level grading connects examination performance to learning outcomes, but manual marking introduces workload and variation between markers. This study evaluates vision-language models (VLMs) for handwritten outcome-based assessment across five dimensions: accuracy, human agreement, repeated-run reliability, error concentration, and explanation quality. Using 1,982 criterion-level records from 485 undergraduate examination answers, we compare 20 configurations spanning Qwen2.5-VL, InternVL3, Pixtral, a Donut baseline, and a cascade ensemble. Evaluation setups include zero-shot prompting, few-shot prompting, partial fine-tuning, and Low-Rank Adaptation (LoRA). Two independent faculty markers regraded all 291 test criteria, providing a human agreement baseline on the same assessment materials. Qwen2.5-VL with LoRA achieved Quadratic Weighted Kappa (QWK) of 0.727 and mean absolute error of 0.435 marks against the examiner, compared with mean human-pair QWK of 0.551. This comparison reflects calibration to the examiner's training marks. LoRA outperformed partial fine-tuning for all three instruction-tuned VLMs, while few-shot prompting reduced QWK in every configuration with valid prompted scores. Aggregate reliability and exact repeatability diverged: intraclass correlations ranged from 0.790 to 0.874, yet 50.2-63.6% of criteria changed marks across five sampled runs. Attention-guided deletion showed no statistically significant advantage over random masking, and four faculty reviewers reached no consensus on explanation usefulness. These findings highlight the need for rubric-specific calibration, repeatable scoring, review of consequential errors, and separate validation of explanations. The released evaluation protocol supports criterion-level assessment research and grading tools with teacher oversight.

↑