智能体排行榜的测量优先审计:污染敏感性、匹配对照再评估与评分器验证
Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation
浏览论文内容
中文总结 AI 辅助
提出测量优先审计框架,区分智能体排行榜污染渠道,通过匹配对照实验发现基准相关差距证据不足,强调需来源证据与验证评分器。
中文摘要 AI 辅助
智能体排行榜越来越多地在公共基准上评估系统,而这些基准的任务陈述和包含解决方案的工件可能仍然可访问。我们提出了一种测量优先的审计框架,根据支持污染声明所需的证据来区分这些声明。该框架区分了需要不同证据的三个渠道:训练时暴露、评估时检索以及流水线/脚手架泄漏。在故障关闭规则下,每个渠道被编码为开放、部分、封闭或未知。在九个Holistic Agent Leaderboard(HAL)配置中,27项渠道评估均未被编码为封闭,但在四个配置中确认了事件。随后,我们将该框架的行为组件应用于SWE-bench Verified上报告的文件定位差距,采用结果盲、同仓库匹配对照设计,并进行对称提示泄漏筛选、配对及仓库感知的不确定性分析以及评分器验证,在GPT-4.1和DeepSeek-V4-Flash上评估。在对称筛选和配对完整性排除后保留的100对中,GPT-4.1显示出+10.0分的配对加权Top-3基准相关差距,但预设配对自助程序和事后仓库平衡分析的95%置信区间均包含零,使得基准相关差距尚无定论。复现评分器未通过其验证门槛:针对共识人工标签,两种模型均无法建立足够的评分器敏感性,且DeepSeek-V4-Flash在正确金标准比较上的两次触发均为假阳性。在没有来源证据、适当对照、对称泄漏筛选和经过验证的评分器的情况下,更强的污染声明是不合理的。结果并未确证训练数据成员资格、污染普遍性或基准导致的分数膨胀。
英文摘要
Agentic leaderboards increasingly evaluate systems on public benchmarks whose task statements and solution-bearing artifacts can remain accessible. We propose a measurement-first audit framework that distinguishes contamination claims according to the evidence required to support them. It separates three channels that require different evidence: training-time exposure, evaluation-time retrieval, and pipeline/scaffold leakage. Each channel is coded as open, partial, closed, or unknown under a fail-closed rule. Across nine Holistic Agent Leaderboard (HAL) configurations, none of the 27 channel assessments was coded closed, but incidents were confirmed in four configurations. We then apply the behavioral component of the framework to a reported file-localization gap on SWE-bench Verified, using an outcome-blind, same-repository matched-control design with symmetric prompt-leakage screening, paired and repository-aware uncertainty analyses, and scorer validation, evaluated on GPT-4.1 and DeepSeek-V4-Flash. Among the 100 pairs retained after symmetric screening and the pair-integrity exclusion, GPT-4.1 showed a $+10.0$-point pair-weighted Top-3 benchmark-associated gap, but the 95\% intervals from both the prespecified paired-bootstrap procedure and the post-hoc repository-balanced analysis included zero, leaving the benchmark-associated gap inconclusive. The reproduction scorer did not pass its validation gate: against consensus human labels, sufficient scorer sensitivity could not be established for either model, and both DeepSeek-V4-Flash firings on correct-gold comparisons were false positives. Without provenance evidence, appropriate controls, symmetric leakage screening, and validated scorers, stronger contamination claims are not warranted. The results do not establish training-data membership, contamination prevalence, or benchmark-induced score inflation.
发表机构
- Northeastern University(东北大学)
- The George Washington University(乔治华盛顿大学)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。