超越重叠:估计基准暴露的因果效应
LeakScale: Estimating the Causal Effect of Benchmark Exposure
浏览论文内容
中文总结 AI 辅助
本文提出LeakScale干预框架,通过创建需私有信息的可执行任务估计基准暴露的因果效应,在多个模型与领域上量化了暴露带来的准确率提升。
中文摘要 AI 辅助
有证据表明,评估材料进入训练集并不能揭示其对评估的影响程度。这一区别使得受污染的基准分数难以解读:来源可以确定接触,但只有反事实才能量化归因于该接触的性能。我们提出了LeakScale,一个用于估计这一缺失量的干预性框架。LeakScale创建全新的可执行任务,这些任务需要公共任务中不存在且无法推导的私有、特定于家族的信息,控制对该信息的访问,并估计由此产生的控制调整后的可执行准确率变化。在2,048个独特家族、两个模型家族、两个可执行领域和262,144次生成中,暴露在每个模型与领域的组合中都提升了准确率,增益范围从+7.17到+27.31个百分点。这些发现区分了两个常被混淆的实证问题:基准接触是否发生,以及报告分数在多大程度上依赖于该接触。LeakScale使后者可直接测量。
英文摘要
Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact. We present LeakScale, an interventional framework for estimating this missing quantity. LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-derivable from the public task, controls access to that information, and estimates the resulting control-adjusted change in executable accuracy. Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improves accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points. These findings separate two empirical questions that are often conflated: whether benchmark contact occurred and how strongly a reported score depends on it. LeakScale makes the latter directly measurable.
发表机构
- University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。