增强视觉异常检测的可解释性自动化
Augmenting Visual Anomaly Detection with Automated Interpretability
- Politecnico di Milano(米兰理工大学)
- Thales Alenia Space(泰雷兹阿莱尼亚宇航公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究利用稀疏自编码器和多模态大语言模型自动分解并干预异常检测信号,抑制干扰、放大异常特征,显著提升PatchCore在多个基准上的检测性能。
AI中文摘要:
视觉异常检测器识别与已知正常数据的偏差,但其异常信号可能混合了真实异常的证据与良性视觉变化。我们研究自动化可解释性是否能够通过识别并干预该信号的不同组成部分来增强视觉异常检测器。我们使用稀疏自编码器(SAEs)将PatchCore最近邻残差分解为稀疏特征,并向多模态大语言模型(MLLM)提供高激活和对比性非激活示例,该模型描述每个特征并将其标记为异常、干扰或不确定。这些标签指导SAE隐藏表示中的干预,其中干扰特征被抑制,异常特征被放大。编辑后的表示随后用于重建补丁嵌入,并使用PatchCore重新评分。在来自四个基准的40个类别中,联合应用两种干预将源数据上的宏平均图像级AUROC从0.8724提高到0.8857,在合成损坏下从0.8066提高到0.8210。在具有真实采集偏移的三个额外RobustAD类别中,相同的干预将源数据上的AUROC从0.8745提高到0.9056,在真实采集偏移下从0.6069提高到0.6599。最后,在所有43个类别中的单个特征干预表明,MLLM标签在总体上与特征对正常和异常图像的不同影响方式一致。
英文摘要:
Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different components of this signal. We decompose PatchCore nearest-normal residuals into sparse features using Sparse Autoencoders (SAEs), and provide high-activation and contrastive non-active examples to a Multimodal LLM, which describes each feature and labels it as anomaly, distractor, or uncertain. These labels guide interventions in the SAE hidden representation, where distractor features are suppressed and anomaly features amplified. The edited representation is then used to reconstruct patch embeddings, which are rescored with PatchCore. Across 40 categories from four benchmarks, applying both interventions jointly improves macro-average image-level AUROC from 0.8724 to 0.8857 on source data and from 0.8066 to 0.8210 under synthetic corruptions. On three additional RobustAD categories with real acquisition shifts, the same interventions improve AUROC from 0.8745 to 0.9056 on source data and from 0.6069 to 0.6599 under real acquisition shifts. Finally, individual feature interventions across all 43 categories show that the MLLM labels are aligned in aggregate with how features differently affect normal and anomalous images.