GeoMAD:基于可变形融合与分布对齐的几何感知多视图异常检测
GeoMAD: Geometry-Aware Multi-View Anomaly Detection via Deformable Fusion and Distributional Alignment
- National Taiwan University(台湾大学)
- National Taiwan Normal University(台湾师范大学)
- Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究院)
- Microsoft Taiwan Corporation(微软台湾公司)
- VinUniversity
- National Taiwan University of Science and Technology(台湾科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
GeoMAD是一种多视图多类别异常检测框架,通过跨视图可变形融合模块与分布视图对齐,实现几何感知与分布一致的高效融合,在Real-IAD等数据集上表现出色。
AI中文摘要:
多视图异常检测(MvAD)通过利用多个相机视角的互补观测来检测缺陷,核心挑战是在融合视角时具备充分的几何感知能力,同时能扩展至多类别工业场景。现有方法通常分为两个极端:基于体素的融合提供显式几何对齐,但需要昂贵的3D构建和类别特定假设;而轻量的基于块的融合效率高,但依赖离散候选匹配,缺乏连续的跨视图对应关系。本文提出GeoMAD,一个统一的多视图、多类别AD框架,以解决几何对应不足和分布不一致问题。我们的跨视图可变形融合模块(CDFM)直接在2D特征图上学习内容自适应、视图对特定的采样偏移,并将其排列在多尺度窗口金字塔中,结合图像全局参考采样,无需相机标定、体素构建或类别特定的3D监督即可实现分层跨视图对应。我们进一步引入分布视图对齐(DVA),一种自监督跨视图正则化损失,将每个视图的瓶颈分布与按实例的视图中心目标对齐,无需像素级对应即可强制全局一致性。CDFM和DVA结合,桥接了局部几何对应与全局分布一致性,在保持2D特征空间学习效率的同时,提供几何感知和分布一致的融合。在Real-IAD和MANTA-Tiny上的大量实验表明,GeoMAD在统一MvAD中实现了出色的检测和定位性能。
英文摘要:
Multi-view anomaly detection (MvAD) detects defects by exploiting complementary observations from multiple camera viewpoints. The central challenge is to fuse views with sufficient geometric awareness while remaining scalable to multi-class industrial settings. Existing methods typically fall into two extremes: voxel-based fusion provides explicit geometric alignment but requires costly 3D construction and class-specific assumptions, whereas lightweight patch-based fusion is efficient but relies on discrete candidate matching and lacks continuous cross-view correspondence. In this paper, we propose GeoMAD, a unified multi-view, multi-class AD framework that addresses both geometric correspondence deficiency and distributional inconsistency. Our \textit{Cross-view Deformable Fusion Module} (CDFM) learns content-adaptive, view-pair-specific sampling offsets directly on 2D feature maps and arranges them across a multi-scale window pyramid with image-global reference sampling, enabling hierarchical cross-view correspondence without camera calibration, voxel construction, or class-specific 3D supervision. We further introduce \textit{Distributional View Alignment} (DVA), a self-supervised cross-view regularization loss that aligns each view's bottleneck distribution against a per-instance view-centric target, enforcing global consistency without pixel-level correspondence. Together, CDFM and DVA bridge local geometric correspondence and global distributional consistency, providing geometry-aware and distribution-consistent fusion while preserving the efficiency of 2D feature-space learning. Extensive experiments on Real-IAD and MANTA-Tiny show that GeoMAD achieves strong detection and localization performance in unified MvAD.