重新思考多模态零样本异常检测中的辅助模态:从语义融合到条件调制
Rethinking Auxiliary Modalities in Multimodal Zero-shot Anomaly Detection: From Semantic Fusion to Conditional Modulation
AI总结:
本研究针对现有多模态零样本异常检测方法的缺陷,提出即插即用的辅助条件增强框架,通过全局到局部的条件调制实现选择性多模态增强,在MVTec 3D-AD等数据集上提升了现有RGB零样本异常检测器的性能并达到最优。
AI中文摘要:
近期,基于基础模型的方法通过视觉-语言预训练为RGB图像赋予了强大的零样本异常检测(ZSAD)能力。然而,仅靠RGB观测在感知以几何变形、深度变化或细微表面变化为主的异常时仍存在局限。辅助模态可提供互补的结构信息,但现有的多模态方法通常将其直接融合到共享语义空间中,这可能会干扰RGB基础模型建立的文本对齐异常语义,且往往需要特定模态的架构。为解决该问题,我们提出一种即插即用的辅助条件增强框架用于零样本异常检测。该框架不重构联合多模态异常语义空间,而是保留原始RGB图像-文本异常匹配通路,并将辅助观测作为条件信号用于RGB特征细化,使辅助模态能无缝增强现有的基于RGB的零样本异常检测器。具体而言,一个轻量级元学习模块以全局RGB和辅助表示为输入,生成样本自适应的低秩残差更新,以确定应如何细化RGB特征;我们还从初始RGB异常响应和辅助可靠性中构建了感知不确定性的空间调制,用于确定应增强或抑制局部残差更新的位置。这种从全局到局部的条件调制可在保留原始RGB异常语义的同时实现选择性多模态增强。在MVTec 3D-AD和Eyecandies上的大量实验表明,我们的框架能持续提升多个流行的基于RGB的零样本异常检测器的性能,实现了多模态零样本异常检测的最优性能。
英文摘要:
Recent foundation model-based methods have endowed RGB images with strong zero-shot anomaly detection (ZSAD) through vision-language pretraining. However, RGB observations alone remain limited in perceiving anomalies dominated by geometric deformation, depth variation, or subtle surface changes. Auxiliary modalities can provide complementary structural information, but existing multimodal methods typically fuse them directly into a shared semantic space, which may disturb the text-aligned anomaly semantics established by RGB foundation models and often requires modality-specific architectures. To address this issue, we propose a plug-and-play auxiliary-conditioned enhancement framework for zero-shot anomaly detection. Instead of reconstructing a joint multimodal anomaly semantic space, our framework preserves the original RGB image-text anomaly matching pathway and uses auxiliary observations as conditional signals for RGB feature refinement, allowing auxiliary modalities to seamlessly enhance existing RGB-based zero-shot anomaly detectors. Specifically, a lightweight meta-learning module takes global RGB and auxiliary representations as input and generates sample-adaptive low-rank residual updates to determine how RGB features should be refined. We further construct uncertainty-aware spatial modulation from the initial RGB anomaly response and auxiliary reliability, which determines where local residual updates are strengthened or suppressed. This global-to-local conditional modulation enables selective multimodal enhancement while preserving the original RGB anomaly semantics. Extensive experiments on MVTec 3D-AD and Eyecandies demonstrate that our framework consistently improves multiple popular RGB-based zero-shot anomaly detectors, achieving state-of-the-art performance for multimodal zero-shot anomaly detection.