arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据集身份,而非新颖性:OOD检测增益虚高的来源

Dataset Identity, Not Novelty: The Source of an Inflated OOD Detection Gain

Donghoon Lee, Shinjin Kang

arXiv 2610.01096首次发表:更新:

发表机构

Hongik University(弘益大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示OOD检测中报告的增益主要源于数据集身份识别而非新颖性,通过留出整个OOD数据集测量膨胀,并提供闭式解以预测哪些拟合会虚高。

AI 中文摘要

事后(post-hoc)分布外(OOD)检测器读取训练好的分类器的激活值并返回一个分数。该分数在分布内数据上拟合,而评估它的基准测试为拟合本身提供第二份OOD数据。一些检测器在其上调整一个常数;另一些则在特征空间中拟合一个方向,或训练一个灵活的合并器,并将其达到的数字报告为仍可获得的增益。每一次这样的拟合都在同一OOD数据集的留出样本上进行验证。这种检查排除了对单个图像的记忆,但并未排除拟合转而学习其正在查看的是哪个数据集的情况;识别一个OOD数据集而非新颖性的方向能完美通过验证。从业者安装的检测器会遇到来自未拟合来源的OOD数据,因此这一差异决定了所报告数字的价值。我们通过留出整个OOD数据集(而非其样本)来测量该差异,并称这一差距为“膨胀”。我们在ImageNet和CIFAR-100骨干网络上跨一系列合并器读取该值。通常协议报告的大部分增益实际上是数据集身份而非新颖性。拟合的大小并不影响存留的部分,因此该效应并非普通过拟合。其占比反而取决于输入是否暴露类别身份,而两个仅改变该属性的控制实验在两种基准的每个骨干网络上都能分离出膨胀。一个闭式解解释了该效应,并从拟合行计算它,因此从业者无需运行留出协议即可判断哪些拟合会膨胀。这些拟合中有一个存留,即该领域已在指定验证数据集上选择的单一常数。任何高于它的拟合所报告的增益都是留出协议无法返回的,而在一个基准上,存留的部分下降,而报告的部分上升。

英文摘要

A post-hoc out-of-distribution (OOD) detector reads the activations of a trained classifier and returns a score. It fits that score on in-distribution data, and the benchmarks that evaluate it supply a second piece of OOD data for the fitting itself. Some detectors tune a constant on it. Others fit a direction in feature space or train a flexible combiner and report the number that it reaches as the gain that is still available. Every such fit is validated on held-out samples of the same OOD dataset. That check rules out memorizing individual images. It says nothing about a fit that has instead learned which dataset it is looking at, and a direction that recognizes one OOD dataset rather than novelty passes it perfectly. The detector that a practitioner installs meets OOD data from a source that nobody fitted it on, so the difference decides what the reported number is worth. We measure it by holding out the whole OOD dataset rather than a sample of it, and we call that gap the inflation. We read it across a range of combiners on ImageNet and CIFAR-100 backbones. Most of the gain that the usual protocol reports turns out to be dataset identity rather than novelty. The size of the fit does not move what survives, so the effect is not ordinary overfitting. The share depends instead on whether the input exposes class identity, and two controls that vary that property alone separate the inflation on every backbone of both benchmarks. A closed form accounts for the effect and computes it from the fitting rows, so a practitioner can tell which fits will inflate without running the hold-out protocol. One of these fits survives, namely the single constant that the field already picks on a designated validation dataset. Anything above it reports a gain that the hold-out protocol does not return, and on one benchmark what survives falls while what is reported climbs.

Comments30 pages, 7 figures, 33 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑