arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29887cs.CL

查看计分板:多项选择题评估的计分方案分析

Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

  • University of Maryland(马里兰大学)
  • New York University(纽约大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Nishant Balepur, Paiheng Xu, Wei Ai, Eunsol Choi, Rachel Rudinger, Jordan Boyd-Graber

AI总结:

该研究分析6种教育启发式计分方案对MCQA评估的影响,发现其可改变31个LLM的排名、更准确预测用户偏好,并揭示不同模型能力,还探讨了方案向其他任务的扩展。

AI中文摘要:

NLP中的多项选择题问答(MCQA)基准采用答对计分(准确率),但在教育测试中,计分方案(模型遵循的答题模式组合与答题评分规则)是决定奖励何种能力的关键设计选择。我们采用6种受教育启发的计分方案(评估准确率之外的能力:干扰项排除、弃权(不执行)、置信度校准与自我修正),探究答对计分的替代方案如何改变MCQA的测量内容。在大语言模型(LLM)基准上,这些方案:1)使31个LLM的排名在重新表述的答对提示之外发生变化;2)更准确预测LLM Arena中用户偏好的LLM;3)揭示不同模型能力,如GPT-5极少弃权(不执行)且易于自我修正,而较弱的开放权重模型常弃权(不执行)且不愿排除选项。鉴于替代计分方案的益处,我们讨论将其扩展到多项选择题之外任务的方法。

英文摘要:

Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.

补充信息

↑