Fi-ImageNet-1k:来自ImageNet-1k验证集内部的OOD基准
Fi-ImageNet-1k: An OOD Benchmark From the Inside of the ImageNet-1k Validation Set
- Czech Technical University in Prague(布拉格捷克技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究构建了更具挑战性的OOD基准Fi-ImageNet-1k,利用ImageNet-1k验证集的标注错误生成,其FPR@95挑战性较NINCO提升3.8倍,现有方法难以达到低假阳性率。
AI中文摘要:
分布外(OOD)检测用于预测测试图像是否不属于任何预定义类别。为评估该任务,基准需要来自分布内(ID)数据之外的图像,通常这些图像是临时定义或收集的。由于没有完美的真值,带有ID标签的数据集本身就包含了OOD图像的自然来源。我们利用这类标注错误,提出了Fi-ImageNet-1k,这是一个基于ImageNet-1k验证集图像构建的OOD数据集,这些图像被近期的ReImageNet重新标注工作归为不属于任何ImageNet-1k类别。每张图像都在MLLMs、VLMs和反向图像搜索提供的证据支持下,由专业人类标注人员检查,并与所有视觉相似的ID类别进行比较。我们仅保留那些可被分配到ImageNet-1k标签空间之外特定类别的图像。最终得到的Fi-ImageNet-1k包含522个类别的655张图像,比任何常用的OOD数据集都更具挑战性。在95%真阳性率(TPR)下,所评估的分类器与OOD检测器的任何组合都无法达到低于51%的假阳性率(FPR@95)。与近期的NINCO相比,我们的数据集对于最先进的监督式OOD检测方法,在FPR@95指标上的挑战性提升了3.8倍。
英文摘要:
Out-of-distribution (OOD) detection predicts whether a test image belongs to none of the predefined classes. To evaluate this task, benchmarks need images from outside the in-distribution (ID) data; typically, these are defined or collected in an ad hoc fashion. Since no ground truth is perfect, ID-labeled datasets themselves contain a natural source of OOD images. We exploit such annotation errors and present Fi-ImageNet-1k, an OOD dataset built from ImageNet-1k validation images that the recent ReImageNet reannotation effort assigned to no ImageNet-1k class. Each image was examined by expert human annotators supported by evidence from MLLMs, VLMs, and reverse image search, comparing it against all visually similar ID classes. We keep only images that could be assigned a specific class outside the ImageNet-1k label space. The resulting Fi-ImageNet-1k, with 655 images from 522 classes, is substantially more challenging than any commonly used OOD dataset. No evaluated combination of classifier and OOD detector achieves a false positive rate below 51% at 95% true positive rate (FPR@95). Compared to the recent NINCO, our dataset is 3.8x more challenging in the FPR@95 metric for state-of-the-art supervised OOD detection methods.