发表机构
MVTec Software GmbH; Technical University of Munich; Amazon; Intel; NVIDIA(MVTec软件有限公司; 慕尼黑工业大学; 亚马逊; 英特尔; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有异常检测基准不切实际的问题,提出VAND 4.0应用驱动基准,覆盖工业制造与零售物流,发现无监督分割仍有挑战、VLM难替代专用模型,并引入效率度量与罕见缺陷数据集。
AI 中文摘要
现有的异常检测基准已趋于饱和且往往不切实际。作为 VAND 4.0 挑战赛的一部分,我们引入了一个隐藏测试、应用驱动的基准,涵盖两个部署关键领域:工业制造和零售物流。在工业赛道(MVTec AD 2)中,结果显示无监督异常分割仍然具有挑战性:最佳常规设置方法仅达到约 57% 的像素级 $SegF_1$,表明仍有很大的改进空间。零样本方法落后约 15 个 $SegF_1$ 点,证实了针对正常数据的任务特定训练对于精确缺陷定位仍然至关重要。对分布偏移的鲁棒性仍然是一个关键开放挑战,DINOv3 骨干网络明显主导该赛道。在零售赛道(Kaputt 2)中,结果显示:(1) 对于常见缺陷类型,监督缺陷检测已接近饱和;(2) 最佳现成 VLM 方法落后于专门模型约 28 AP,证实目前 VLM 无法替代微调检测器;(3) 参考图像对顶级方法并未证明有帮助。在罕见缺陷上性能崩溃(溢出约 53 AP,缺失单元约 27 AP),其中监督上限受数据可用性限制。为了推动该领域的未来进展,我们提供了一个新的低患病率零售异常检测数据集(Kaputt-Rare)。在两个赛道中,计算效率作为一等指标进行评估,结合了性能、吞吐量、内存和功耗。我们引入了一种新的效率度量指标,并揭示顶级方法依赖重型架构,而效率在很大程度上被忽视。总体而言,我们得出结论,社区需要 (a) 更多效率感知的方法开发,以及 (b) 针对罕见缺陷和变化条件的真正异常检测方法。此 https URL。
英文摘要
Existing Anomaly Detection benchmarks are saturated and often unrealistic. As part of the VAND 4.0 Challenge, we introduce a hidden-test, application-driven benchmark across two deployment-critical domains: industrial manufacturing and retail logistics. In the Industrial Track (MVTec AD 2), the results reveal that unsupervised anomaly segmentation remains challenging: the best regular-setting method achieves only ~57\% pixel-level $SegF_1$, indicating substantial room for improvement. Zero-shot approaches trail by ~15 $SegF_1$ points, confirming that task-specific training on normal data remains essential for precise defect localization. Robustness to distribution shifts remains a key open challenge and DINOv3-backbones clearly dominate this track. In the Retail Track (Kaputt 2), the results reveal that (1) supervised defect detection is approaching saturation for common defect types; (2) the best off-the-shelf VLM approach trails specialized models by ~28 AP, confirming that currently VLMs cannot replace fine-tuned detectors, (3) reference images did not prove helpful for top-performing approaches. Performance collapses on rare defects (spillage ~53 AP, missing units ~27 AP), where the supervised ceiling is bounded by data availability. To drive future progress in this domain, we provide a new low-prevalence retail AD dataset (Kaputt-Rare). Across both tracks, computational efficiency is assessed as a first-class metric combining performance, throughput, memory, and power consumption. We introduce a novel metric for measuring efficiency and reveal that that top-performing methods rely on heavy architectures while efficiency is largely neglected. Overall, we conclude that the community needs (a) more efficiency-aware method development, and (b) true anomaly detection approaches for rare defects and shifting conditions. https://sites.google.com/view/vand4-cvpr2026/challenge