arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05743cs.CV

ConceptADapt:用于少样本工业异常检测的概念引导自适应特征重建与动态注意力机制

ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection

Yufei Li, Yicheng Ruan, Long Tian, Dongsheng Wang, Liang Bao

AI总结:

针对少样本工业异常检测中正常样本稀缺导致的泛化性问题,提出ConceptADapt模型,通过概念引导特征重建与动态注意力机制,在三个基准上优于现有最优方法。

AI中文摘要:

少样本工业异常检测(FS-IAD)聚焦于冷启动阶段工业检测中视觉缺陷的检测与定位,每类仅提供有限的正常训练样本。该领域近期进展主要利用基础模型的视觉特征,已取得可观性能。尽管基础模型特征具备强表征能力,但因正常训练样本极度稀缺,模型泛化性仍较脆弱。为解决这一关键问题,我们提出ConceptADapt,一种结合动态注意力的概念引导自适应特征重建模型。具体而言,该模型从有限的支持特征中预学习一组固定的正常概念,并利用这些概念挖掘与查询特征的关系,从而在测试时重新校准特征统计以提升异常检测性能。为缓解低数据 regime 下尤为严重的特征捷径问题,我们进一步开发了与稀疏自编码器集成的动态注意力机制,用于在训练阶段学习鲁棒的正常概念。此外,为实现推理阶段的快速适配,我们通过在注意力模块中引入LoRA保持模型轻量化,仅引入极少的更新参数。在MVTec-AD、VisA和MPDD三个广泛采用的FS-IAD基准上开展的实验表明,我们的模型在检测和定位任务中均持续优于当前最优(SOTA)方法,且在各类 shot 设置下均实现了显著提升。

英文摘要:

Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase, where only a limited number of normal training samples are available per category. Recent advances in this field predominantly leverage visual features from foundation-model and have achieved promising performance. Despite the strong representational power of foundation-model features, the model generalization remains fragile due to the extreme scarcity of normal training data.To address this pivotal issue, we propose ConceptADapt, a concept-guided adaptive feature reconstruction model with dynamic attention. Specifically, our model pre-learns a set of fixed normal concepts from the limited support features and leverages them to mine relationships with query features, thereby recalibrating their statistics for improved anomaly detection at test time. To mitigate the prevalent feature shortcut problem, which is particularly severe under low-data regimes, we further develop a dynamic attention mechanism integrated with sparse autoencoders to learn robust normal concepts during training. Moreover, to enable fast adaptation during inference, our model remains lightweight by incorporating LoRA into the attention module, which introduces only minimal updating parameters.Extensive experiments on three widely adopted FS-IAD benchmarks, including MVTec-AD, VisA, and MPDD, demonstrate that our model consistently outperforms state-of-the-art (SOTA) approaches across both detection and localization tasks, achieving significant improvements under various shot settings.

补充信息

↑