arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17938cs.CLcs.AI

评分需要评分标准,而非智能

Grading Needs a Rubric, Not Intelligence

  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Jhen-Ke Lin

AI总结:

该研究提出any-to-bench设计原则,证实小型语言模型依据明确评分标准评分的可靠性与高成本模型相当,评分标准中的官方答案是关键,可将评分与评分者智能解耦。

AI中文摘要:

当小型语言模型依据明确的评分标准评分时,其对开放式考试答案的评分可靠性可与成本高得多的模型相媲美。我们将这一主张作为 any-to-bench 的设计原则进行测试:前沿模型在数据摄入阶段读取一次源文档,以提取每个问题及其评分标准;随后成本较低的模型完成所有重复的评分工作。我们评估了来自两个模型家族、三个推理努力水平下的六种高性价比模型配置。每种配置回答24道开放式考试问题,且每种配置还对每份答卷进行三次评分,共产生每道题3456个评分。评分结果主要取决于被评分的答案:答案身份解释了95.6%的评分方差,而评分者身份仅解释了0.2%。提高答题者的推理努力可使获得的分数变动多达满分的0.143,而提高评分者的推理努力最多仅使分配的分数变动0.006。作为验证,我们添加了六个前沿级评分者,其评分结果与前述一致,且作为一组时可靠性并未更高。随后的两项消融实验在相同问题和答案上对评分标准进行分解:移除其标准和等级但保留官方答案,未产生可测量的变化;同时移除官方答案则会使可靠性崩溃(组内相关系数ICC从0.888降至0.628)、分数膨胀,且使评分者的推理努力再次产生影响。评分标准是将评分与评分者智能解耦的关键,而在评分标准中,官方答案承担了几乎全部的作用。我们未发现采用评分标准锚定评分时存在长度偏好或同家族偏好的证据。

英文摘要:

Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.

↑