arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25168cs.CVcs.MM

看得更多,检测得更少?抑制多视图异常检测中的信息泄漏

See More, Detect Less? Taming Information Leakage in Multi-View Anomaly Detection

  • National Taiwan University(台湾大学)
  • Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室)
  • National Taiwan Normal University(台湾师范大学)
  • VinUniversity
  • Microsoft Taiwan Corporation(微软台湾公司)
  • National Taiwan University of Science and Technology(台湾科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Shang-Fu Chen, Kuan-Chuan Peng, Jhih-Ciang Wu, Wen-Huang Cheng, Kai-Lung Hua

AI总结:

该研究针对多视图异常检测中的跨视图信息泄漏问题,提出GLAD框架,结合视觉基础模型特征与MMA、OGA模块,在Real-IAD等数据集上优于现有最优方法,证实信息限制的重要性。

AI中文摘要:

在多视图异常检测中,更多的跨视图信息实际上可能会产生负面影响。当在基于重构的流程中简单融合多个检测视图时,来自完整视图的正常线索会传播到解码器,解码器会忠实地重构异常区域,从而破坏检测器所依赖的重构差距。我们将这种失效模式称为“跨视图信息泄漏”,并证明有效的多视图融合必须明确限制到达解码器的信息。基于这一见解,我们提出了GLAD(Global-Local Attention Driven framework,全局-局部注意力驱动框架),这是首个将视觉基础模型特征与局部和全局跨视图融合相结合的多视图异常检测框架。多视图融合注意力(MMA)模块以线性复杂度执行局部跨视图融合,具备可学习的视图重要性加权和令牌级门控,使每个视图能以O(N)的成本有选择地整合其他视图的细粒度证据。对象引导注意力(OGA)模块通过将所有视图的类别令牌聚合为单个对象级表示,并通过温度缩放的sigmoid门控将其广播回补丁令牌,从而捕获全局上下文,该模块会替换原始补丁表示而非添加残差,以保留重构差距。在Real-IAD和MANTA-Tiny数据集上的实验表明,GLAD在样本级、图像级和像素级指标上均优于现有最优方法,证实了合理的信息限制是多视图异常推理的关键。

英文摘要:

In multi-view anomaly detection, more cross-view information can actually hurt. When multiple inspection views are naively fused in a reconstruction-based pipeline, normal cues from intact views propagate to the decoder, which faithfully reconstructs anomalous regions, collapsing the reconstruction gap the detector depends on. We call this failure mode \emph{cross-view information leakage} and show that effective multi-view fusion must explicitly restrict the information reaching the decoder. Building on this insight, we present GLAD(Global-Local Attention Driven framework), the first framework combining vision foundation model features with local and global cross-view fusion for multi-view anomaly detection. The Multi-view Merging Attention (MMA) module performs local cross-view fusion at linear complexity with learnable view importance weighting and token-wise gating, letting each view selectively incorporate fine-grained evidence from other views at $\mathcal{O}(N)$ cost. The Object-Guided Attention (OGA) module captures global context by aggregating class tokens from all views into a single object-level representation and broadcasting it back to patch tokens via temperature-scaled sigmoid gating, replacing the original patch representations rather than adding a residual to preserve the reconstruction gap. Experiments on Real-IAD and MANTA-Tiny show that GLAD outperforms state-of-the-art methods across sample-, image-, and pixel-level metrics, confirming that principled information restriction is key to multi-view anomaly reasoning.

↑