过度可分性:用于基准污染检测的干扰控制残差流探测
Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection
浏览论文内容
中文总结 AI 辅助
该研究针对基准污染检测的现有方法缺陷,提出了基于残差流探测的干扰控制方案,通过水平匹配基线等设计降低假阳性率,在真实Transformer上验证了方案有效性并发布了相关实现。
中文摘要 AI 辅助
当前基准污染检测采用n元重叠、基于似然的成员推理或金丝雀字符串方法,每种方法均需通常不可获得的资源:训练语料、精心选择的检验统计量或数据集发布时的预见性。近期一种替代方法通过对内部激活的线性探测读取污染情况。我们表明,该方法的自然实现方式并不奏效,并提出了一种可通过测量的方案。该方案报告探测准确率深度轮廓上的零和对比,该轮廓以水平匹配的安慰剂基线为中心,针对标签置换原假设进行检验,参考集大小为可疑集的两倍。每个选择均替换了我们测量并拒绝的更简单替代方案。报告过度可分性的水平而非其形状,使得假阳性率可追踪分析师自身控制集的大小,在真实原假设下从0.03到0.99变化。针对平坦深度轮廓的对比在两个方向均失效:当表面可解码性随深度上升时,72%的情况会拒绝真实原假设;当表面可解码性下降时,则完全丧失功效。项目自助法固定拟合的探测,在置换原假设(其会重新拟合探测)下拒绝至多9%的情况,而后者仅拒绝2%的情况。基线集大小减半会使错误率增至三倍。在真实Transformer上,基线深度轮廓明显不平坦,在时间划分上跨度达29.1个准确率点,其非平坦性追踪项目集之间的表面差异(6次审计的相关系数为0.87),因此校正恰好在最需要的地方最大。所有4个匹配的Pile分支均返回原假设,该方案对时间划分拒绝作出判定而非给出结果。这并未确立Transformer是否具有熟悉度方向:唯一的正结果出现在交换性失效的划分上。实现、测试和审计已发布。
英文摘要
Benchmark contamination is diagnosed with n-gram overlap, likelihood-based membership inference, or canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at release. A recent alternative reads it off a linear probe on internal activations. We show the natural way to do this does not work, specify one that survives measurement, then find that the correction making it work carries more variance than the null it is tested against. The protocol reports a zero-sum contrast on the depth profile of probe accuracy, recentred on a level-matched placebo baseline, tested against a label-permutation null, with the reference set twice the size of the suspect set. Each choice replaces a simpler alternative we rejected on measurement. Reporting the level of excess separability rather than its shape makes the false positive rate track the size of the analyst's own control set, 0.03 to 0.99 under a true null. Contrasting against a flat depth profile rejects a true null 0.72 of the time when surface decodability rises with depth, and loses all power when it falls. On real transformers the protocol fails a test the simulations did not pose. The recentring subtracts an estimate, and the permutation null holds it fixed. Re-estimated across split seeds on four audits of contaminated checkpoints, its standard deviation is 1.30 to 1.56 times the null's own in every arm: what is subtracted to remove a bias is more variable than what it corrects. The one nominally significant result, p = 0.0075, becomes 0.0745 once that variance is propagated, and no verdict is issued. The simulations missed this because their surface key is the covariate driving item variation; on real text it is a proxy, and degrading key quality in simulation reproduces it. We add a companion measurement and a widened null. No arm shows contamination.