arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00986cs.CV

从费舍尔信息视角理解和克服多模态异常检测中的跨模态融合偏差

Understanding and Overcoming Cross-modal Fusion Bias in Multimodal Anomaly Detection From A Fisher Information Perspective

Kaifang Long, Lianbo Ma, Liming Liu, Guoyang Xie

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态异常检测中被忽视的跨模态融合偏差问题,提出即插即用框架UCFB,经实验验证在多种设置下均实现性能提升。

中文摘要 AI 辅助

当前多模态异常检测(Multimodal Anomaly Detection, MAD)的进展主要依赖于增强多模态融合,尤其是通过整合RGB和深度(Depth)数据来获取更丰富的异常表示。然而,跨模态融合偏差这一多模态学习中众所周知的挑战,在MAD领域却较少受到关注。这一空白催生了一个关键问题:我们能否克服这种偏差以突破当前工作的性能瓶颈?在本文中,我们首先通过费舍尔信息矩阵分析跨模态融合偏差对MAD的影响。随后,基于这些发现,我们提出了UCFB——一个简单却有效的即插即用框架,旨在缓解MAD中的跨模态融合偏差。该框架通过联合采用费舍尔信息引导的动态校准(以调整模态特定的正则化权重)和典型相似度分析(以改善模态间交互)来实现这一目标。在MVTec 3D-AD和Eyecandies数据集上进行的大量实验表明,UCFB在单类别、多类别和少样本设置中均取得了持续的性能提升。

英文摘要

Current advancements in Multimodal Anomaly Detection (MAD) are largely driven by enhancing multimodal fusion, particularly through the integration of RGB and Depth data for richer anomaly representation. However, less attention was devoted to analyzing the role of cross-modal fusion bias, a well-known challenge in multimodal learning, in MAD. This gap motivates a key question: can we overcome this bias to break the performance bottleneck of current work? In this paper, we first analyze the impact of cross-modal fusion bias in MAD via the Fisher Information Matrix. Then, grounded in these findings, we propose UCFB, a simple yet effective plug-and-play framework designed to mitigate cross-modal fusion bias in MAD. It achieves this by jointly employing Fisher-information-guided dynamic calibration to adjust modality-specific regularization weights and canonical similarity analysis to improve inter-modal interactions. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets demonstrate that UCFB achieves consistent improvements in single-class, multi-class, and few-shot settings.

发表机构

  • Northeastern University(东北大学)
  • CATL(宁德时代)

机构由 AI 辅助整理,请以论文原文为准。

↑