发表机构
Universidad Europea de Madrid; Universidad Europea de Valencia(欧洲马德里大学; 欧洲瓦伦西亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种按被挫败的缓解措施分类的基准污染体系,实施四字段披露协议,对 41 份文档测量工具可靠性,发现无文档涵盖全部五种污染类型,为基准污染研究提供分类框架与测量方案。
AI 中文摘要
基准得分是模型、评估工具、引出预算、抽样总体以及污染状态的共同属性。排行榜公布模型与得分,因此能力与泄漏在观测上等价。现有分类体系针对自动检测对污染进行分类,而非研究人员在发表时面临的问题:在已应用缓解措施的情况下,哪些有效性威胁仍未解决?我们引入一种按每种类型所挫败的缓解措施组织的分类体系——直接型、衍生型、时间型、分布型和获得型——涵盖训练时与评估时的泄漏。保留私人测试集仅能单独解决第一种类型;第五种类型在评估过程中获得,由于它是单次运行的属性,必须随报告的得分而非基准发布一同记录。我们将其实施为四字段披露协议,其中“未知”是有效条目,以 CC BY 4.0 许可发布,附带 JSON Schema、验证器和示例。设计团队外的两名编码人员使用预注册工具对 41 份文档进行应用。在 29 份主流程文档上,各变量的线性加权 κ 值介于 0.00 至 0.35 之间(中位数为 0.21),而单编码人员重测上限为 0.84,因类别偏斜导致结果崩溃;合并后通过偶然校正将其提升至 0.46,而非达成更好的一致性。两个变量低于预先注册的流行度稳健阈值:分层报告与本文引入的获得型。分歧集中在变量何时适用,而非文档陈述内容。13% 的文档报告了引出预算,无文档涵盖全部五种类型。本研究的贡献是该分类体系、由此产生的得分侧制品,以及对工具可靠性与当前披露情况的预注册测量。
英文摘要
A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status. Leaderboards publish the model and the score, so capability and leakage stay observationally equivalent. Existing taxonomies classify contamination for automated detection, not the question a reporter faces at publication: given the mitigations already applied, which validity threats remain open? We introduce a taxonomy organized by the mitigation each type defeats -- direct, derivative, temporal, distributional, and acquired -- spanning training-time and evaluation-time leakage. Holding out a private test set closes the first alone. The fifth is acquired during the evaluation itself; because it is a property of one run, it must be recorded with the reported score rather than with the benchmark release. We operationalize it as a four-field disclosure protocol in which "unknown" is a valid entry, released under CC BY 4.0 with a JSON Schema, a validator, and worked examples. Two coders external to the design team applied a pre-registered instrument to 41 documents. Per-variable linear-weighted $κ$ runs from 0.00 to 0.35 (median 0.21) over 29 main-pass documents against a single-coder test-retest ceiling of 0.84, collapsing under the class skew the registration anticipated; pooling raises it to 0.46 through chance correction rather than better agreement. Two variables fall below the prevalence-robust threshold registered in advance: strata reporting and the acquired type introduced here. Disagreement concentrates on when a variable applies rather than on what a document states. Elicitation budgets are reported in 13% of documents, and no document addresses all five types. The contribution is the taxonomy, the score-side artifact that follows from it, and a pre-registered measurement of instrument reliability and current disclosure.
Comments18 pages, 3 figures, 6 tables. Specification, audit instrument, pre-registration and analysis code: https://github.com/Jangulo7/contamination-disclosure-paper