arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04210cs.CV

PADFormer:基于稀疏视角图像的姿态无关异常检测

PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images

Ruiqi Wang, Yiming Qian, Fenggen Yu, Yuxuan Lu, Dakuo Wang, Hao Zhang, Jing Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有姿态无关异常检测方法依赖3D重建的缺陷,提出PADFormer模型,利用ViT结合跨视图掩码重建等机制,在PAD基准及FSAD任务上实现最优性能,兼具效率与泛化性。

中文摘要 AI 辅助

姿态无关异常检测(Pose-agnostic Anomaly Detection,PAD)仍是一项具有挑战性的任务,因为异常可出现在任意视角下,要求方法能处理显著的姿态变化。现有方法依赖复杂的3D重建,计算成本高且需要大量多视图数据。我们提出PADFormer,一种新颖的图像空间方法,利用视觉Transformer(Vision Transformer,ViT)直接重建查询图像的无异常版本,同时保留姿态信息。我们的核心见解是:通过仅在正常数据上训练,将跨视图掩码重建适配到异常检测中,结合动态补丁选择和空间对齐机制,能在显著姿态变化下从稀疏参考视图中有效学习。推理阶段,我们执行多次带不同掩码模式的前向传播,生成一组无异常重建结果,确保全面覆盖查询图像;通过将这些重建结果与查询图像比较来检测异常。PADFormer在PAD基准上达到了最优结果,同时在经典少样本异常检测(Few-shot Anomaly Detection,FSAD)任务上保持可比性能,展现出无需3D重建的卓越效率和泛化能力。

英文摘要

Pose-agnostic Anomaly Detection (PAD) remains challenging as anomalies can appear under arbitrary viewpoints, requiring methods to handle significant pose variations. Existing approaches rely on complex 3D reconstruction, which are computationally expensive and require extensive multi-view data. We propose PADFormer, a novel image-space approach that leverages Vision Transformer (ViT) to directly reconstruct anomaly-free versions of query images while preserving pose information. Our key insight is to adapt cross-view masked reconstruction for anomaly detection through training exclusively on normal data, combined with dynamic patch selection and spatial alignment mechanisms that enable effective learning from sparse reference views under significant pose variations. During inference, we perform multiple forward passes with different masking patterns to generate an ensemble of anomaly-free reconstructions, ensuring comprehensive coverage of the query image. Anomalies are detected by comparing these reconstructions with the query image. PADFormer achieves state-of-the-art results on the PAD benchmark while maintaining comparable performance on classic few-shot anomaly detection (FSAD) tasks, demonstrating superior efficiency and generalization without requiring 3D reconstruction.

发表机构

  • Amazon(亚马逊)
  • Simon Fraser University(西蒙菲莎大学)
  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑