发表机构
School of Software Engineering, Xi’an Jiaotong University; Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学软件学院; 西安交通大学人工智能与机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对零样本异常检测中正常与异常文本原型语义重叠的问题,提出Proximity-CLIP框架,通过视觉校准语义间隔和异常查询模块,在多个基准上以最小架构改动超越现有方法。
AI 中文摘要
视觉-语言模型为零样本异常检测(ZSAD)提供了一种有前景的方法。然而,由于以物体为中心的偏差,正常和异常的文本原型表现出高度的语义重叠。虽然强制它们之间严格正交可以提高判别性,但将高度连续的视觉输入映射到剧烈正交的原型上会引入几何困境,破坏预训练的结构连续性。为了解决这个问题,我们提出了Proximity-CLIP,一个通过视觉校准语义间隔来引导视觉适应的框架。首先,我们引入了一种视觉校准的语义邻近学习机制,该机制使用有界动态正则化来学习适当的语义间隔,确保判别性分离同时保持结构对齐。其次,我们设计了一个由这些文本先验驱动的异常查询模块(AQM)。利用校准后的异常原型作为语义查询,AQM主动从上下文视觉补丁中检索局部缺陷线索,减轻了全局池化过程中细微异常的稀释。大量实验表明,Proximity-CLIP在多个ZSAD基准上以最小的架构修改超越了当前最先进的方法。
英文摘要
Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to object-centric bias, normal and anomalous text prototypes exhibit a high semantic overlap. While enforcing strict orthogonality between them improves discriminability, mapping highly contiguous visual inputs onto drastically orthogonal prototypes introduces a geometric dilemma, disrupting the pre-trained structural continuity. To address this problem, we propose Proximity-CLIP, a framework that visually calibrates the semantic margin to guide visual adaptation. First, we introduce a visually-calibrated semantic proximity learning mechanism that uses a bounded dynamic regularization to learn an appropriate semantic margin, ensuring discriminative separation while preserving structural alignment. Second, we design an Anomaly Query Module (AQM) driven by these text priors. Using the calibrated anomalous prototype as a semantic query, the AQM actively retrieves localized defect cues from contextual visual patches, mitigating the dilution of subtle anomalies during global pooling. Extensive experiments demonstrate that Proximity-CLIP outperforms current state-of-the-art methods across multiple ZSAD benchmarks with minimal architectural modifications.
CommentsAccepted to ECCV 2026. Including supplementary material