arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可解码但不可检测:用于近离群点基准测试的泄漏指纹

Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks

Vishnu Bindu Balachandran

arXiv 2607.19393首次发表:更新:

AI 中文总结

研究针对基于扰动的离群点检测器在文档基准测试中AUROC低的问题,通过提炼泄漏指纹进行验证,修正协议使扰动信号可解码但不可检测,贡献了修正协议和泄漏诊断,而非新的离群点方法。

AI 中文摘要

在对基于扰动的离群点检测器进行文档基准测试时,我们记录到其曲线下面积(AUROC)为0.326,远低于0.5的随机水平。原因是基准测试泄漏:指定的“离群点”类别是模型训练过的,其示例位于分布内拟合集中,检测器将它们正确排序为熟悉的会受到惩罚。删除该类别并在两个域中重新训练35个模型后,分数提高到0.911。我们将这种污染提炼成一个泄漏指纹——近乎完美的监督可解码性(AUROC约为1)与低于0.65的无监督检测能力相结合——并在52种设置(20种泄漏、32种干净)的受控测试组上进行验证,在嵌入空间中实现了18/20的灵敏度和31/32的特异性;匹配的拟合集排除控制在20/20时完美。对24个标准近/远离群点基准测试对进行的实际测试仅在一个(本质上困难的CIFAR - 100对CIFAR - 10对)上触发,而在远离群点对上未触发,证实了特异性且标准跨数据集构建是干净的。在修正后的协议下,扰动信号是可解码但不可检测的:有监督的读取器可以恢复离群点信号(AUROC为0.87 - 1.00),而无监督检测器则不能,并且扰动方法在普通马氏距离上没有改进。我们从理论上解释了原因,并为了透明性撤回了早期的循环相关性。贡献在于修正后的协议和经过验证的泄漏诊断,而非新的离群点方法。

英文摘要

While auditing a perturbation-based OOD detector on a document benchmark, we recorded an AUROC of 0.326 -- well below the 0.5 chance level. The cause is a benchmark leak: the designated "OOD" class is one the model was trained on, so its examples sit inside the in-distribution fit set and the detector is penalized for correctly ranking them as familiar. Deleting the class and retraining 35 models across two domains raises the score to 0.911. We distill the contamination into a leak fingerprint -- near-perfect supervised decodability (AUROC approximately 1) coupled with unsupervised detection collapsed below 0.65 -- and validate it on a controlled battery of 52 settings (20 leaked, 32 clean) across ResNet-50 and ViT-B/16 on CIFAR-10/100, achieving sensitivity 18/20 and specificity 31/32 in embedding space; the matched fit-set-exclusion controls are perfect at 20/20. An in-the-wild audit of 24 standard near/far OOD benchmark pairs fires on exactly one (the intrinsically hard CIFAR-100 vs CIFAR-10 pair) and on no far-OOD pair, confirming specificity and that standard cross-dataset construction is clean. Under the corrected protocol, perturbation signals are decodable but not detectable: a supervised reader recovers the OOD signal (AUROC 0.87-1.00) while no unsupervised detector does, and the perturbation method does not improve on plain Mahalanobis distance. We provide a theoretical account of why and, for transparency, retract an earlier circular correlation. The contributions are a corrected protocol and a validated leak diagnostic, not a new OOD method.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑