arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

XMatchAD:基于重建的异常检测的跨模态匹配视角

XMatchAD: A Cross-Modal Matching Perspective on Reconstruction-based Anomaly Detection

Mingxiu Cai, Zhe Zhang, Gaochang Wu, Tianyou Chai

arXiv 2607.23658首次发表:更新:

发表机构

Northeastern University; State Key Laboratory of Synthetical Automation for Process Industries(东北大学; 流程工业综合自动化国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基于重建的无监督异常检测方法难以捕捉细微异常、异常边界模糊的问题,提出XMatchAD框架,从伪跨模态匹配视角利用输入与重建图像的匹配关系,通过多步骤提升异常检测和定位精度,性能优于现有方法。

AI 中文摘要

基于重建的方法在无监督异常检测(UAD)中取得显著成功,能通过建模输入图像与其重建图像之间的差异来识别和定位异常。然而,这些方法难以捕捉细微异常,且异常边界模糊,限制了其在复杂多类场景中的有效性。为此,我们提出XMatchAD,一种从伪跨模态匹配视角重新诠释任务的新型UAD框架。具体而言,将输入图像和重建图像视为两个互补模态,利用其匹配关系进行异常检测。首先,使用预训练特征提取器编码判别性表示;其次,引入注意力引导的跨模态匹配机制匹配局部模态间异常相关模式并相互细化特征,提高对不同形状和细微偏差异常的敏感性,提升异常检测和定位精度;第三,设计自适应频率感知融合模块,通过跨模态多尺度表示的高频分量耦合进一步勾勒清晰的异常边界。在MVTec-AD、VisA和MPDD基准上的综合评估表明,我们的方法始终实现卓越性能,在多类异常检测和定位方面优于现有方法。代码将在该https URL发布。

英文摘要

The remarkable success of reconstruction-based methods in Unsupervised Anomaly Detection (UAD) lies in their ability to identify and localize anomalies by modeling discrepancies between input images and their reconstructed counterparts. However, these approaches often struggle to capture subtle anomalies and tend to produce blurred anomaly boundaries, which significantly limits their effectiveness, particularly in complex multi-class scenarios. To address these issues, we present XMatchAD, a novel UAD framework that reinterprets the task from a pseudo cross-modal matching perspective. Specifically, the input and reconstructed images are treated as two complementary modalities and their matching relationships are precisely exploited for anomaly detection. First, a pre-trained feature extractor is employed to encode discriminative representations. Second, an attention-guided cross-modal matching mechanism is introduced to match local inter-modal anomaly-related patterns while mutually refining the features. This enhances the sensitivity to anomalies with diverse shapes and subtle deviations and significantly improves the precision of anomaly detection and localization. Third, we design an adaptive frequency-aware fusion module that further delineates sharp anomaly boundaries through the coupling of high-frequency components from cross-modal multi-scale representations. Comprehensive evaluations on MVTec-AD, VisA, and MPDD benchmarks demonstrate that our method consistently achieves superior performance, outperforming state-of-the-art methods in multi-class anomaly detection and localization. The code will be released at https://github.com/Mingxiu-Cai/XMatchAD.

Commentsaccepted by IEEE Transactions on Image Processing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑