AI 中文总结
该研究提出敏感性网格,发现评估框架的配置差异使LLM分数呈区间变化,配置敏感项目承载模型间大部分分数差,且框架可决定排行榜排名,压缩方法保留此类项目。
AI 中文摘要
多项选择题基准测试固定了问题和正确答案,但未固定评估框架(harness):选项的顺序、提示词的措辞,以及语言模型的答案是从生成文本中读取还是从每个选项的似然中读取。关于该评估框架敏感性的研究将其报告为 aggregate 分数方差,却未研究方差落在哪些项目上,以及这些项目是否是区分不同模型的关键。我们将大语言模型(LLM)的评估框架视为自变量,并将其影响分解到单个项目上。我们引入了「敏感性网格(fragility grid)」:来自4个系列的12个开放权重指令调优LLM,在26种同等合理的评估框架配置下回答来自4个基准测试(ARC、HellaSwag、MMLU、TruthfulQA)的相同3679个项目,为每个模型、项目和配置记录一个正确性比特。该比较是匹配的,因为项目、权重和贪婪解码保持固定,仅评估框架发生变化。在该网格下,模型的分数是一个区间而非一个点:gemma4-31b的分数仅因评估框架不同,在31%至89%之间。由此得出三个结果:在两个相邻模型均稳定回答的项目上,二者得分持平,且对配置敏感的项目平均承载了一对模型之间95.7%的分数差距;12个模型中有4个在某些配置下排名第一,即评估框架决定了优胜者;基准压缩方法最大化的项目区分度与敏感性的相关性为0.28(95%置信区间为0.25至0.30),因此压缩过程保留了敏感项目而非将其移除。决定分数的是评分选择,而非协议通常固定的选项顺序。我们发布了项目级记录和分析脚本,可在CPU上几秒内重新生成所有数字,并将敏感性网格定位为排行榜报告排名前可运行的检查项。
英文摘要
Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next. We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items. We introduce the \textit{fragility grid}: 12 open-weight instruction-tuned LLMs from 4 families answer the same 3{,}679 items from 4 benchmarks (ARC, HellaSwag, MMLU, TruthfulQA) under 26 equally defensible harness configurations, recording one correctness bit for every model, item, and configuration. The comparison is matched, since the items, the weights, and the greedy decoding stay fixed while only the harness varies. Under the grid a model's score is a band rather than a point: gemma4-31b scores between 31 and 89 percent depending only on the harness. Three results follow. On the items that two adjacent models both answer stably the pair is tied, and config-fragile items carry 95.7 percent of a pair's gap on average. Four of the 12 models reach rank one under some configuration, so the harness selects the winner. Item discrimination, the property that benchmark-compression methods maximize, correlates with fragility at 0.28 (95 percent CI 0.25 to 0.30), so compression keeps the fragile items rather than removing them. The scoring choice, not the option order that protocols usually fix, is the load-bearing axis. We release the per-item records and the analysis script, from which every number regenerates on a CPU in seconds, and we position the fragility grid as a check a leaderboard can run before it reports an order.
Comments25 pages in total