位置偏差隐藏在天花板效应背后:LLM基准测试的排列诊断
Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks
浏览论文内容
中文总结 AI 辅助
研究LLM基准测试中位置偏差问题,引入inspect_permute工具对四个供应商在五个MMLU主题上进行测试,发现位置偏差仅在特定准确率区间可检测,明确了可检测区域及两种机制类型,界定了位置偏差测量范围,使核心问题可验证。
中文摘要 AI 辅助
在多项选择题的LLM评估中,位置偏差被广泛认为是能力比较中的一个混淆因素,但已发表的测量依赖于单一答案顺序的洗牌,其结果将偏差信号与内容级噪声和采样随机性混淆。本文引入了inspect_permute,这是inspect_ai评估框架的开源扩展,它对每个问题进行详尽的答案顺序排列,并报告位置偏差的卡方/Cramer V特征以及自举置信区间。作者在五个MMLU主题上对四个供应商(gpt-4o-mini、claude-haiku-4-5、gemini-2.5-flash、grok-3)应用了该工具,在温度为0的生成下进行了24,000次API调用,并在观察到一半数据之前通过公共SHA-256哈希预先注册了伪造预测。结果表明,位置偏差仅在大约60-95%的基本准确率的“金发姑娘区”内具有统计可检测性。低于该区域,处理负载占主导地位,淹没了特定主题的信号;高于该区域,天花板效应将方差压缩到卡方检验分辨率以下。可检测的单元格分为两种机制类型:单调的A到D下降(处理负载,在低层级模型中)和非单调的D下降(内容模糊性,在狭窄的能力范围内)。标准MMLU将每个前沿层级模型置于检测带之上,因此在那里没有信号应被理解为不可测量,而不是无偏差。与arXiv:2606.26185中的天花板效应特征一起,这项工作界定了位置偏差测量的可检测区域,并使该领域的核心问题能够以可验证的形式提出。该软件包、数据和预注册遵循麻省理工学院许可。
英文摘要
Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single answer-order shuffles whose results confound the bias signal with content-level noise and sampling stochasticity. I introduce inspect_permute, an open-source extension to the inspect_ai evaluation framework that runs exhaustive answer-order permutations per question and reports the chi-squared / Cramer V signature of position bias with bootstrap confidence intervals. I apply the tool across four vendors (gpt-4o-mini, claude-haiku-4-5, gemini-2.5-flash, grok-3) on five MMLU subjects, 24,000 API calls under temperature-0 generation, with falsifier predictions pre-registered via a public SHA-256 hash before half the data was observed. Position bias turns out to be statistically detectable only within a roughly 60-95% base-accuracy Goldilocks zone. Below it, processing-load dominance swamps subject-specific signal; above it, ceiling effects compress the variance below the chi-squared test resolution. Detectable cells separate into two mechanism types: monotone A-to-D decrease (processing_load, in low-tier models) and non-monotone D-drop (content_ambiguity, in a narrow capability band). Standard MMLU places every frontier-tier model above the detection band, so absence of signal there should be read as not measurable, not unbiased. Together with the ceiling-effect characterisation in arXiv:2606.26185, this work brackets the detectable region of position-bias measurement and makes the field central question askable in a verifiable form. Package, data, preregistration under MIT.