TED:面向提示异常定位的文本轴证据分解
TED:Text-Axis Evidence Decomposition for Prompted Anomaly Localization
浏览论文内容
中文总结 AI 辅助
提出TED事后评分方法,通过文本轴证据分解区分缺陷与困难正常区域,无需训练即可显著提升CLIP异常定位性能,平均增益达+10.9。
中文摘要 AI 辅助
CLIP是一个强大的视觉-语言模型,但其并非为细粒度缺陷定位而设计;因此,基于CLIP的异常检测器通过提示或轻量模块对其进行适配,以提高缺陷敏感性。我们表明,更强的敏感性并不必然使局部证据可靠:在领域偏移下,适配后的CLIP-AD模型常常对真实缺陷和视觉上复杂的正常区域均赋予高异常分数。问题并非简单地缺失缺陷信息,而是局部评分规则将缺陷与困难正常证据解码为相同的异常证据。我们提出TED(文本轴证据分解),一种事后评分方法,用于判断每个模糊响应是更受源缺陷补丁支持,还是更受被误判为异常的源正常补丁支持。TED在宿主模型的正常-对-异常文本响应下比较这些支持度,保持骨干网络和提示不变,且无需目标域训练。它可作为原始VLM骨干的无训练分数,或作为适配后CLIP-AD宿主的源校准残差修正。在冻结的VLM骨干上,TED在像素级定位上显著优于原始提示相似度;在适配后的宿主上,它在大多数像素级设置中优于P-AUROC、P-PRO和P-AP。在更强的困难假阳性竞争下增益最大,平均定位增益从低竞争场景的+5.0提升至中/高竞争场景的约+10.9。这些结果表明,可恢复的缺陷证据已存在于预训练多模态表示中,但可靠定位需针对困难正常竞争者进行解码。代码将在TED GitHub仓库发布。
英文摘要
CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions. The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-normal evidence, having the same anomaly evidence. We propose TED (Text-Axis Evidence Decomposition), a post-hoc scoring method that asks whether each ambiguous response is better supported by source defect patches or by source normal patches mistaken as anomalous. TED compares these supports under the host's normal-versus-anomaly text response, leaves the backbone and prompts unchanged, and requires no target-domain training. It works as a train-free score for raw VLM backbones or as a source-calibrated residual correction for adapted CLIP-AD hosts. Across frozen VLM backbones, TED substantially improves pixel-level localization over raw prompt similarity; across adapted hosts, it improves most pixel-level settings over P-AUROC, P-PRO, and P-AP. Gains are largest under stronger hard-FP competition, with mean localization gain increasing from +5.0 in low-competition regimes to about +10.9 in mid/high-competition regimes. These results suggest that recoverable defect evidence can already exist in pretrained multimodal representations, but reliable localization requires decoding it against hard-normal competitors. Code will be released at TED GitHub repository.
发表机构
- Chung-Ang University(中央大学)
- SNUAILAB Co., Ltd.(SNUAILAB 有限公司)
机构由 AI 辅助整理,请以论文原文为准。