UniScale:任意尺度的工业异常生成
UniScale: Arbitrary-Scale Industrial Anomaly Generation
- School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对工业异常检测中真实异常样本不足及小尺度异常易丢失的问题,提出UniScale框架,通过EMT策略和生成后融合去噪方法,在VisA、MVTec AD 2等数据集上显著优于现有方法。
AI中文摘要:
工业异常检测面临的一个主要挑战是缺乏真实世界的异常样本。虽然生成模型被用于创建异常数据,但现有方法在处理小尺度异常时仍然存在困难,这一问题的出现是因为扩散模型中的极端下采样会导致小异常的信息在潜在空间中丢失。为解决这一问题,我们引入了UniScale,一个用于生成高保真工业异常的统一训练和推理框架,支持任意尺度的异常生成。在训练阶段,我们提出了误差抑制多尺度训练(EMT)策略,该策略使模型能够学习到异常丰富的位置感知纹理,同时抑制纹理获取中上采样引起的插值误差,确保模型既能学习小尺度异常,又能对常规尺度异常保持有效。在推理阶段,我们提出了“生成后融合去噪”策略,将异常生成与背景集成解耦,防止小异常被模糊。实验表明,我们的方法在异常生成质量和下游检测任务上均优于当前最先进的竞争对手:在VisA数据集上实现了相对IS(a)提升45.86%(从1.81提升至2.64),在MVTec AD 2数据集上实现相对IS(a)提升37.70%(从1.22提升至1.68);同时在VisA数据集上将下游像素级IoU提升4.22%,在MVTec AD 2数据集上将AUROC提升6.55%。代码可在this https URL获取。
英文摘要:
Industrial anomaly inspection faces a major challenge due to the lack of real-world anomaly samples. While generative models are used to create anomaly data, existing methods still struggle when handling small-scale anomalies. This failure occurs because extreme downsampling in diffusion models causes the information of small anomalies to be lost in the latent space. To address this, we introduce UniScale, a unified training and inference framework for high-fidelity industrial anomaly generation across arbitrary scales. During training, we introduce an Error-Suppressed Multi-Scale Training (EMT) strategy, which enables the model to learn the rich location-aware textures of anomalies, while suppressing upsampling-induced interpolation errors in texture acquisition, ensuring the model is capable of learning small-scale anomalies, while remaining effective for regular scale anomalies. For inference, we propose Generation-then-Fusion Denoising. It decouples anomaly generation from background integration, preventing small anomalies from being overwhelmed. Extensive experiments demonstrate that our method outperforms state-of-the-art competitors in both anomaly generation quality and downstream detection performance. It achieves a relative IS(a) improvement of 45.86% (from 1.81 to 2.64) on VisA and 37.70% (from 1.22 to 1.68) on MVTec AD 2, while also improving the downstream pixel-level IoU by 4.22% on VisA and AUROC by 6.55% on MVTec AD 2. Code is available at https://github.com/HUST-SLOW/UniScale.