查看计分板:多项选择题评估的计分方案分析
Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation
- University of Maryland(马里兰大学)
- New York University(纽约大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究分析6种教育启发式计分方案对MCQA评估的影响,发现其可改变31个LLM的排名、更准确预测用户偏好,并揭示不同模型能力,还探讨了方案向其他任务的扩展。
AI中文摘要:
NLP中的多项选择题问答(MCQA)基准采用答对计分(准确率),但在教育测试中,计分方案(模型遵循的答题模式组合与答题评分规则)是决定奖励何种能力的关键设计选择。我们采用6种受教育启发的计分方案(评估准确率之外的能力:干扰项排除、弃权(不执行)、置信度校准与自我修正),探究答对计分的替代方案如何改变MCQA的测量内容。在大语言模型(LLM)基准上,这些方案:1)使31个LLM的排名在重新表述的答对提示之外发生变化;2)更准确预测LLM Arena中用户偏好的LLM;3)揭示不同模型能力,如GPT-5极少弃权(不执行)且易于自我修正,而较弱的开放权重模型常弃权(不执行)且不愿排除选项。鉴于替代计分方案的益处,我们讨论将其扩展到多项选择题之外任务的方法。
英文摘要:
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.