双锚点,效果更佳:用于零样本异常检测的分层分组合并方法
Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection
浏览论文内容
中文总结 AI 辅助
针对现有CLIP-based零样本异常检测方法的局限,提出Dual-Anchor框架,结合分层图像锚点与文本锚点,在8个工业和6个医疗基准上实现了优异泛化性能。
中文摘要 AI 辅助
零样本异常检测(ZSAD)旨在识别未见过领域中的异常,该设置对存在域偏移的工业和医疗应用尤为关键。然而,大多数基于CLIP的ZSAD方法仅将语义锚定在文本模态上,导致性能对提示设计高度敏感且视觉定位能力薄弱。为缓解这些局限,我们提出了Dual-Anchor框架,该框架通过自上而下的分组机制构建分层图像锚点,以补充传统文本锚点。此机制逐步聚合从局部到全局的图像特征,形成正常和异常组令牌,这些令牌作为图像锚点,并在Group-Gated Token Refiner中充当门控信号,以增强全局表示。随后将细化的图像锚点与文本提示融合,构建动态状态提示。通过联合强化视觉和文本语义,我们的框架稳定了图像-文本对齐,降低了对提示的依赖,并在8个工业和6个医疗基准上实现了出色的泛化能力。
英文摘要
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen domains, a setting that is particularly critical for industrial and medical applications where domain shifts are prevalent. However, most CLIP-based ZSAD methods anchor semantics solely on the text modality, making performance highly sensitive to prompt design and leading to weak visual grounding. To mitigate these limitations, we propose a Dual-Anchor framework that complements conventional text anchors with hierarchical image anchors constructed via a top-down grouping mechanism. This mechanism progressively aggregates local-to-global image features to form normal and abnormal group tokens, which serve as image anchors and act as gating signals in a Group-Gated Token Refiner to enhance the global representation. The refined image anchors are then fused with text prompts to construct dynamic state prompts. By jointly reinforcing visual and textual semantics, our framework stabilizes image-text alignment, reduces prompt dependency, and achieves strong generalization across 8 industrial and 6 medical benchmarks.
发表机构
- Sogang University(西江大学)
- LG Electronics(LG电子)
- NAVER Cloud(NAVER云)
机构由 AI 辅助整理,请以论文原文为准。