arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DuoAD:利用[CLS]双重特征进行无训练少样本异常检测

DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection

Jyun-Ze Tang, Po-Han Huang, Ming-Ching Chang, Chih-Fan Hsu, Jeng-Lin Li

arXiv 2607.23924首次发表:更新:

发表机构

Inventec Corporation; University at Albany, State University of New York(英业达股份有限公司; 纽约州立大学奥尔巴尼分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有视觉基础模型无训练异常检测依赖局部特征、未充分利用全局信息的问题,提出DuoAD框架,利用ViT [CLS]令牌双重特征,通过自动增强选择和注意力引导特征重加权实现无训练少样本异常检测,取得优异结果并建立新的技术水平。

AI 中文摘要

视觉基础模型实现了强大的无训练异常检测。然而,现有方法大多依赖独立局部补丁特征,未充分利用视觉变换器(ViT)编码的全局上下文信息。本文识别出ViT [CLS]令牌的双重特征,其嵌入提供异常不变全局语义表示,注意力图隐含突出空间异常区域。基于此,提出全自动异常检测框架,引入由[CLS]级语义一致性驱动的自动增强选择策略和注意力引导特征重新加权机制。该方法在多级别特征上集成这些组件,无需训练或参数调整即可实现稳定异常评分和精确定位。在单样本设置下,在MVTec-AD、VisA和Real-IAD上分别达到97.7%、93.2%和84.5%的Image-AUC分数。该方法建立了即插即用、无训练异常检测的新的技术水平,同时保持强大的鲁棒性和实际可扩展性。

英文摘要

Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on independent local patch features, leaving the global contextual information encoded by Vision Transformers (ViTs) underexploited. In this work, we identify the dual characteristics of the ViT [CLS] token: its embedding provides anomaly-invariant global semantic representation, while its attention maps implicitly highlight spatially abnormal regions. Building on this observation, we propose a fully automated AD framework leveraging global context to remove manual tunings. Our framework introduces (1) an automatic augmentation selection strategy driven by [CLS]-level semantic consistency, and (2) an attention-guided feature reweighting mechanism that dynamically adjusts patch contributions according to [CLS] attention saliency. By integrating these components over multi-level features, our method achieves stable anomaly scoring and precise localization without training or parameter tuning. Under the one-shot setting, it achieves Image-AUC scores of 97.7%, 93.2%, and 84.5% on MVTec-AD, VisA, and Real-IAD. Using a single fixed configuration across categories, backbones, and datasets, the method establishes a new state-of-the-art for plug-and-play, training-free anomaly detection while maintaining strong robustness and practical scalability.

CommentsCode: https://github.com/inventec-ai-center/DuoAD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑