arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM 评分器在何处成功与失效:来自两门计算机科学考试的证据

Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan

arXiv 2609.29333首次发表:更新:

发表机构

King Abdullah University of Science and Technology (KAUST); Oxford Brookes University; University of Oxford(阿卜杜拉国王科技大学; 牛津布鲁克斯大学; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过两门计算机科学考试的大规模实验,揭示 LLM 评分器在特定提示词下可能崩溃或拒绝评分,并证明轻量 LoRA 微调可修复该问题,使小型开源模型达到人类评分员水平。

AI 中文摘要

大型课程中的一次长文考试需要耗费数百小时的评分员工作量,而合格的评分员十分稀缺;LLM 评分器成为一种诱人的替代方案。为揭示其缺陷,我们在 171 种配置下对一场实用的计算机视觉考试(570 名学生,双人评分)进行评分,配置涵盖闭源和开源权重模型;最佳模型的平均绝对误差达到 1.64/35,低于两名人类评分员彼此之间的 2.61/35。问题在于提示词:一段简短的“严格评分员”前言导致 17 个开源权重模型中的 14 个偏离评分区间(MAE ≥ 8),其中三个模型完全停止评分。损害可追溯至前言中两条扣分语句,而非语气或模型规模;其中一条“绝不给予部分分数”单独就使三个受测模型中的两个停止评分。三家供应商的闭源旗舰模型在该前言下校准发生变化,但仍保持在评分区间内。在另一门课程的第二场独立机器学习考试(1,038 名学生,双人评分)的 162 种进一步配置中,该前言使十个模型表现恶化,其中三个偏离区间陷入崩溃,一个拒绝评分,但改善了七个中性提示词下过度评分的模型:该漏洞可复现,但其方向因考试而异。轻量 LoRA 微调可修复此问题:在两场考试合并的约 3,900 个评分样本上训练的一个适配器,使五个小型开源模型在评分员对一致性方面达到或超过人类评分员水平,且对三种严苛人格的敏感性几乎消失(MAE ≤ 0.32)。我们发布了匿名数据集、完整消融网格以及评分、微调和分析流水线。

英文摘要

One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights models; the best reaches mean absolute error $1.64/35$, below the $2.61/35$ two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives $14$ of $17$ open-weights models out of the graded band ($\text{MAE} \ge 8$), three stopping grading altogether. The damage traces to the preamble's two credit-withholding sentences, not to tone or model scale; one of them, ''never give partial credit'', alone makes two of three probed models stop grading. The closed flagships of three vendors shift calibration under it but stay in the band. In $162$ further configurations on a second, independent Machine Learning exam from another course ($1{,}038$ dual-graded students), the preamble worsens ten models, moving three out of the band into collapse and one into refusal, yet improves seven whose neutral prompts over-mark: the vulnerability replicates, but its direction is exam-specific. Light LoRA fine-tuning repairs it: one adapter on the two exams' pooled $\sim 3{,}900$ graded examples brings five small open models to parity or better with a human grader in agreement with the grader pair, and sensitivity to the three harsh personas nearly vanishes ($\le 0.32$ MAE). We release the anonymised dataset, full ablation grid, and grading, fine-tuning and analysis pipelines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑