DriftAD:用于小样本工业异常检测的视觉引导文本漂移
DriftAD: Visually-Guided Text Drift for Few-Shot Industrial Anomaly Detection
浏览论文内容
中文总结 AI 辅助
针对现有小样本工业异常检测方法无法捕捉缺陷局部性与尺度依赖性的问题,提出含ASA、VGTD、DGSG模块及对应损失的DriftAD框架,在MVTec-AD和VisA数据集的1/2/4次设置下取得最优性能。
中文摘要 AI 辅助
小样本异常检测(FSAD)近期从CLIP等视觉-语言模型中受益,这些模型通过将视觉特征与正常和异常状态的文本描述对齐来实现异常检测。然而,现有方法通常依赖静态文本提示,这些提示在整个特征层级和空间维度上统一应用,这种僵化的全局到局部匹配无法捕捉工业缺陷高度局部化且依赖尺度的物理变化。为解决这一问题,我们提出DriftAD,这是一个构建于三个关键模块之上的FSAD框架。首先,异常信号放大(ASA)模块在文本-视觉匹配前通过空间分支和频率分支增强细微缺陷信号。其次,视觉引导文本漂移(VGTD)动态变换冻结的CLIP文本嵌入,将其调整为每一层编码器深度下以局部视觉上下文为条件的分层、空间自适应异常描述符。第三,漂移引导空间门控(DGSG)利用漂移后的异常描述符作为空间探针,选择性增强与异常相关的视觉特征。此外,漂移分离损失防止漂移描述符的表示崩溃,门控监督损失强制DGSG中具有空间区分性的门控。在MVTec-AD和VisA上的大量实验表明,该方法在1次、2次和4次设置下的图像级和像素级指标上均达到了最先进的性能。代码可在this https URL获取。
英文摘要
Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection by aligning visual features with text descriptions of normal and abnormal states. However, existing methods typically rely on static text prompts that are applied uniformly across the entire feature hierarchy and spatial dimensions. This rigid global-to-local matching fails to capture the highly localized and scale-dependent physical variations of industrial defects. To address this, we propose DriftAD, a FSAD framework built on three key modules. First, an Anomaly Signal Amplification (ASA) module enhances subtle defect signals through spatial and frequency branches before text-visual matching. Second, Visually-Guided Text Drift (VGTD) dynamically transforms frozen CLIP text embeddings, steering them into layer?wise, spatially-adaptive anomaly descriptors conditioned on local visual context at each encoder depth. Third, Drift-Guided Spatial Gating (DGSG) uses the drifted abnormal descriptor as a spatial probe to selectively enhance anomaly-relevant visual features. Addi?tionally, a drift separation loss prevents representational collapse of the drifted descriptors, and a gate supervision loss enforces spatially discriminative gating in DGSG. Extensive experiments on MVTec?AD and VisA demonstrate state-of-the-art performance across all 1-, 2-, and 4-shot settings on both image-level and pixel-level metrics. Code is available at https://github.com/wenyang001/DriftAD.
发表机构
- Nanyang Technological University(南洋理工大学)
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。