AI 中文总结
该研究提出VFAD框架,结合变分语义提示与频率自适应表示学习,在13个工业和医学基准上,其零样本异常检测性能优于现有最先进方法。
AI 中文摘要
零样本异常检测(ZSAD)旨在无需目标特定训练数据的情况下,检测和定位未见类别中的异常。尽管近期基于CLIP的方法通过视觉-语言对齐展现出良好的泛化能力,但在捕捉多样异常语义和细微局部变化方面仍存在局限。为解决这些局限,我们提出VFAD,这是一个结合变分语义提示与频率自适应表示学习的统一框架。具体而言,我们引入变分语义提示提取器(VSPE),其从密集的patch token中自适应聚合与异常相关的局部语义,并通过变分信息瓶颈对其进行正则化,从而融入细粒度视觉线索并实现更精准的跨模态对齐。此外,我们开发频率自适应表示聚合(FARA)模块,该模块利用基于小波的频率分解和特定频率的专家聚合来增强异常判别性视觉表示。通过联合强化语义引导和视觉表示学习,VFAD提升了异常判别能力和细粒度定位效果。在13个工业和医学基准上开展的大量实验表明,VFAD在各类异常场景中均持续优于现有最先进的ZSAD方法,代码将在发表后公开。
英文摘要
Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although recent CLIP-based methods have demonstrated promising generalization through vision-language alignment, they remain limited in capturing diverse anomaly semantics and subtle local variations. To address these limitations, we propose VFAD, a unified framework that combines variational semantic prompting with frequency-adaptive representation learning. Specifically, we introduce a Variational Semantic Prompt Extractor (VSPE), which adaptively aggregates anomaly-relevant local semantics from dense patch tokens and regularizes them through a variational information bottleneck, thereby incorporating fine-grained visual cues and enabling more precise cross-modal alignment. Furthermore, we develop a Frequency-Adaptive Representation Aggregation (FARA) module that leverages wavelet-based frequency decomposition and frequency-specific expert aggregation to enhance anomaly-discriminative visual representations. By jointly strengthening semantic guidance and visual representation learning, VFAD improves both anomaly discrimination and fine-grained localization. Extensive experiments on 13 industrial and medical benchmarks demonstrate that VFAD consistently outperforms existing state-of-the-art ZSAD methods across diverse anomaly scenarios. The code will be publicly available upon publication.