发表机构
School of Electronic and Electrical Engineering, University of Sheffield; Ranplan Wireless Network Design Ltd.; Cambridge AI+ Ltd.(谢菲尔德大学电子与电气工程学院; Ranplan无线网络设计有限公司; 剑桥AI+有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过两项预注册审计发现共享端点上黑盒大语言模型评判器的排名一致性未达阈值,分析了三种失效机制,提出改进方案并强调预注册评估需先测量工具有效性。
AI 中文摘要
语言模型评判器如今管控着训练数据、对生成内容打分并驱动着排行榜。这类评判器作为一种测量工具,依赖着一个极少被明确表述的假设:向同一模型名称发送相同的请求,在未来会得到相同的结果。我们在两项所有阈值均预先固定的预注册研究中对这一假设进行了审计;两项研究均未能通过其工具的有效性验证。在52988次审计的请求尝试中,同窗口重复排名的斯皮尔曼相关系数为0.400,而所需阈值为0.90;字节完全相同的次日重放结果的一致性为0.78,所需阈值为0.99,每次测试的执行记录均处于上限水平。存在三种机制解释这一差距:标签到含义的映射对读数的偏差影响与信号本身一样强烈;候选差距比工具自身的噪声 floor 低七个数量级;以及字节完全相同的输入会返回不同的排名,这种精确排列读数会加剧的噪声。在测试的范围内,无论是指标替换还是采样都无法修复这一问题。预注册的后续研究对该问题进行了限定:在采样的天数上等待并未带来帮助(0.805对0.800,在另外五天中重复出现);切换提供商并未带来帮助(四个提供商处于同一水平,中位数为0.74至0.88,且无法通过它们公开的任何元数据字段预测);在批处理不变内核上自托管仅在服务器空闲时有效;在具有已知差距的构造错误上,读数的区分度跟踪错误类型而非大小。我们将证据提炼为三级快照-身份阶梯、八条设计规则和一份报告清单;一项约为研究调用量2%的试点本可以提前暴露那些无法触及的管控门。所有结果均涉及在共享服务基础设施上的外部测量行为。在共享端点上,模型名称并非一个固定的工具;预注册评估必须在对其施加任何管控门之前先测量其工具的有效性。
英文摘要
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.