发表机构
King Abdullah University of Science and Technology (KAUST); Edge Hill University(阿卜杜拉国王科技大学; 埃奇希尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多项选择基准饱和问题,提出AnswerPool方法,通过合并共享上下文的选项池并要求模型分配答案,降低猜测概率并衡量排除法信用,实验显示对所有模型更难且弃权可评分。
AI 中文摘要
多项选择基准测试评分成本低,但已接近饱和,标准的补救措施——编写更难的问题——既缓慢又需针对每个基准重复进行。一个饱和的基准测试仍包含更难的任务。每个问题的错误选项仅针对该问题编写,因此模型可以通过排除几个选项来得分。我们提出AnswerPool:取共享同一上下文的$N$个问题,将所有选项合并到一个列表中,并要求模型为每个问题分配其答案。无需编写新题目,也无需更改标签。对于五个四选项问题,猜对整组的概率从$10^{-3}$降至$5\ imes10^{-7}$,而能识别自身答案的模型保持其多项选择得分,因此因池化而损失的准确率衡量了格式为排除法提供的信用。从池中删除答案会使问题在精确真实标签下无法回答,因此弃权(不执行)在同一轮中被评分。在八个文本、图像和视频基准测试以及十八个模型中,池化对每个模型都更难,排除法信用对最弱的模型最大,且八个开放权重模型中有七个对87%至100%的不可回答问题进行了回答。
英文摘要
Multiple-choice benchmarks are cheap to grade and are running out of room, and the standard remedy, writing harder items, is slow and repeated for every benchmark. A saturated benchmark still holds a harder task. Each question's wrong options are written for that question alone, so a model can score by eliminating a few options. We propose AnswerPool: take $N$ questions that share a context, pool all their options into one list, and ask the model to assign every question its answer. No item is written and no label changes. The chance of guessing a group right falls from $10^{-3}$ to $5\times10^{-7}$ for five four-option questions, and a model that recognizes its answers keeps its multiple-choice score, so the accuracy lost to pooling measures the credit the format gave for elimination. Deleting answers from the pool makes questions unanswerable with exact ground truth, so abstention is scored in the same pass. Across eight text, image, and video benchmarks and eighteen models, pooling is harder for every model, the elimination credit is largest for the weakest models, and seven of eight open-weight models answer 87 to 100% of unanswerable questions.