arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态异常检测:一项综述

Multi-Modal Anomaly Detection: A Survey

Xudong Mou, Zexin Wu, Chuan Luo, Shiru Chen, Xudong Liu, Chunming Hu, Renyu Yang

arXiv 2608.24937首次发表:更新:

发表机构

School of Computer Science and Engineering, Beihang University; School of Software, Beihang University; Shandong Inspur Intelligent Production Technology Co., Ltd(北京航空航天大学计算机科学与工程学院; 北京航空航天大学软件学院; 山东浪潮智能生产技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该综述从假设驱动视角梳理多模态异常检测(MMAD),划分两种互补方法范式,探讨基础模型对MMAD的重塑,汇总基准与评估协议并指出未来方向。

AI 中文摘要

多模态异常检测(MMAD)用于从异构数据源中检测罕见的异常事件,越来越多地应用于工业检测、网络安全等对安全性和可靠性要求极高的场景。然而,现有文献分散在不同领域和模态组合中,且现有综述通常按架构对方法进行分组,而非按多模态场景中异常的定义与区分方式分组。本文从假设驱动的视角对MMAD展开综述:我们对该问题进行形式化,明确其核心挑战背后的5项固有特征,并将现有工作组织为两种互补范式。第一种是正态性假设方法,通过表示学习、跨模态对齐和知识增强来建模规律;第二种是异常性假设方法,通过粗粒度、结构型和语义型异常注入来优化决策边界。我们还探究了基础模型如何通过可扩展预训练、灵活跨模态迁移及新兴推理能力重塑MMAD。最后,我们汇总了跨领域的代表性基准与评估协议,并指出了鲁棒、自适应、可解释MMAD系统的开放问题与未来方向。

英文摘要

Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-critical applications such as industrial inspection and cybersecurity. Yet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnormality is defined and separated in multi-modal settings. We survey MMAD from an assumption-driven perspective. We formalize the problem, identify five intrinsic characteristics underlying its core challenges, and organize prior work into two complementary paradigms. The first, normality-assumption methods, models regularity via representation learning, cross-modal alignment, and knowledge enhancement. The second, anomaly-assumption methods, sharpens decision boundaries through coarse-grained, structural, and semantic anomaly injection. We also investigate how foundation models are reshaping MMAD through scalable pretraining, flexible cross-modal transfer, and emerging reasoning capabilities. Finally, we compile representative benchmarks and evaluation protocols across domains and highlight open problems and future directions for robust, adaptive, and interpretable MMAD systems.

CommentsAccepted for publication in IEEE Transactions on Big Data

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑