arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基准污染何时可检测?信息极限与功率校准审计

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

arXiv 2608.07914首次发表:更新:

发表机构

Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究量化基准污染可检测性的信息极限,提出基于效能的审计方法,发现小样本下高斯预算失效,两阶段规划器可修复预算并弃权,明确非拒绝结果需结合多门限才可解释。

AI 中文摘要

行为污染检测器返回“无证据”,可能是因为基准干净,也可能是因为审计的检测力不足。我们针对以下基准正式化该区分:其中未知比例为α的样本项在训练期间被见过。通过匹配干净对照和见过对照,行为通道为稀疏混合分布Q_α = (1-α)P₀ + αP₁,精确的二阶矩论证表明,可检测性由α·ρ·√m决定,其中ρ² = χ²(P₁ || P₀)衡量行为可分性。任何标量检测器的效能可简化为ef = |E₁f - E₀f| / √Var₀(f) ≤ ρ,该效能可在审计运行前通过对照估计。独立的样本拆分证书可无分布下限估计α,无需定向假设。我们的实证发现具有两面性:冻结的校准效能可预测保留的功率曲线,在6个精确置换通道上的R²为0.83-0.98,但仅基于效能的高斯预算在其规定的小样本量下校准错误,在9/9的门控通过通道中失效,尽管效能本身可迁移。该失效出现在反演而非校准阶段。预先声明的两阶段规划器模拟部署测试,可修复预算,保持一致保守性,并在其探针不可迁移时弃权(不执行)。该证书在审计规模下有效但无实际意义,5个随机种子的配对注入研究准确恢复了机制排序:逐字>释义>表层,其中仅答案的表观信号由基线漂移解释。我们同时报告审计合约及其失效:仅当与产生非拒绝结果的效能、预算和有效性门限结合时,非拒绝结果才具有可解释性。

英文摘要

Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean and seen controls, the behavioral channel is the sparse mixture Q_alpha = (1 - alpha) P_0 + alpha P_1, and an exact second-moment argument shows that detectability is governed by alpha * rho * sqrt(m), where rho^2 = chi^2(P_1 || P_0) measures behavioral separability. Any scalar detector reduces to its efficacy, ef = |E_1 f - E_0 f| / sqrt(Var_0(f)) <= rho, which can be estimated from controls before the audit is run. A separate sample-split certificate lower-bounds alpha distribution-free, without requiring an orientation assumption. Our empirical finding is two-sided. Frozen calibration efficacy predicts held-out power curves, with R^2 = 0.83-0.98 across six exact-permutation channels, but the efficacy-only Gaussian budget is miscalibrated at the small sample sizes it prescribes, failing in 9/9 gate-passing channels even though efficacy itself transports. The failure is in the inversion, not the calibration. A predeclared two-stage planner that simulates the deployed test repairs the budgets, is uniformly conservative, and abstains when its probe does not transport. The certificate is valid but vacuous at audit scale, and a five-seed paired injection study recovers the mechanism ordering verbatim > paraphrase > surface, in which the apparent answer-only signal is explained by baseline drift. We report the audit contract and its failures together: a non-rejection is interpretable only alongside the efficacy, budget, and validity gates that produced it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑