PL-SCEA:为小样本工业异常检测重构预训练注意力机制
PL-SCEA: Reconfiguring Pretrained Attention for Few-Shot Industrial Anomaly Detection
浏览论文内容
中文总结 AI 辅助
本研究提出PL-SCEA,通过重构冻结视觉基础模型的注意力机制,结合轻量级变分自编码器,在MVTec AD、VisA等数据集的小样本设置下,实现了更优的工业异常检测与定位性能。
中文摘要 AI 辅助
视觉基础模型(Vision Foundation Models, VFMs)为小样本工业异常检测提供可迁移的补丁表示,但其注意力计算通常继承自以语义聚合为核心的预训练目标,这可能存在不匹配:支持语义识别的令牌关系可能无法充分暴露异常定位所需的局部纹理与结构偏差。因此,本研究探究冻结VFMs的注意力计算可被重构为异常检测任务相关组件的假设,以幂律自相关增强注意力(Power-Law Self-Correlation Enhanced Attention, PL-SCEA)实例化该思路,其保留预训练查询-键注意力的语义上下文,同时在上下文化的值特征上构建令牌自适应自相关;正相关过滤与幂律重加权随后突出相对于每个令牌关系背景显著的关系,且不引入额外可训练注意力投影。所得特征由轻量级变分自编码器(lightweight variational autoencoder)建模,提供类别特定正态性的固定大小基于重构的表示。两个阶段发挥互补作用:注意力重构塑造局部关系偏差的表示方式,而基于重构的建模将从学习到的正态性的偏差转换为异常分数。在MVTec AD和VisA数据集上,完整框架在评估的小样本设置中实现了有竞争力的图像级检测和始终强劲的像素级定位;消融实验进一步显示,在测试设置下,PL-SCEA结合VAE或记忆库均可提升定位性能。这些结果支持任务对齐的注意力重构可提升冻结预训练表示的异常定位能力的观点。
英文摘要
Vision Foundation Models (VFMs) provide transferable patch representations for few-shot industrial anomaly detection, but their attention computation is typically inherited from pretraining objectives centered on semantic aggregation. This creates a potential mismatch: token relations that support semantic recognition may not adequately expose the localized texture and structural deviations required for anomaly localization. We therefore investigate the hypothesis that the attention computation of a frozen VFM can be reconfigured as a task-relevant component of anomaly detection. We instantiate this idea with Power-Law Self-Correlation Enhanced Attention (PL-SCEA), which retains the semantic context of pretrained query-key attention while constructing token-adaptive self-correlations over contextualized value features. Positive-correlation filtering and power-law reweighting then emphasize relations that are salient relative to each token's relational background, without introducing additional trainable attention projections. The resulting features are modeled by a lightweight variational autoencoder that provides a fixed-size reconstruction-based representation of category-specific normality. The two stages serve complementary roles: attention reconfiguration shapes how local relational deviations are represented, while reconstruction-based modeling converts deviations from learned normality into anomaly scores. Across MVTec AD and VisA, the complete framework achieves competitive image-level detection and consistently strong pixel-level localization across the evaluated few-shot settings. Ablations further show that PL-SCEA improves localization with either the VAE or a memory bank under the tested setting. These results support the view that task-aligned attention reconfiguration can improve the anomaly-localization capability of frozen pretrained representations.
发表机构
- School of Airspace Science and Engineering, Shandong University(山东大学空天科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。