发表机构
Tsinghua University; Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); Shenzhen University(清华大学; 广东省人工智能与数字经济实验室(深圳); 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何利用大型视觉语言模型进行工业异常检测,提出OPD-IAD框架,通过策略内自蒸馏和语言引导视觉锚定,将语言判断转化为像素级异常图,在相关方法中取得最佳性能。
AI 中文摘要
大型视觉语言模型(LVLMs)最近在工业异常检测(IAD)中展现出强大潜力,能提供图像级异常判断和可解释的缺陷推理。但当前基于LVLM的IAD方法难以从生成的语言判断中生成精确的像素级异常图。本文提出OPD-IAD,一个基于证据特权的密集策略内自蒸馏框架,将特权缺陷证据蒸馏到模型自身的策略内判断轨迹上。还引入语言引导视觉锚定,通过判断重转发将图像和问题重新编码为语义锚点,与密集视觉特征对比生成异常图。实验表明OPD-IAD在基于LVLM的IAD方法中性能最佳。
英文摘要
Large vision-language models (LVLMs) have shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable reasoning. However, reliably translating generated judgments into precise pixel-level localization remains challenging. We propose \textbf{M}ixed-Trajectory \textbf{O}n-\textbf{P}olicy \textbf{D}istillation for Language-Guided Industrial \textbf{A}nomaly Detection (MOPDA), the first framework to introduce on-policy self-distillation into LVLM-based IAD. For judgment learning, \method introduces \textbf{Mixed-Trajectory Supervision}, combining student-generated on-policy trajectories with evidence-conditioned teacher trajectories under a shared token-level distillation objective. Student trajectories preserve supervision on deployment-relevant response paths, while teacher trajectories provide complementary evidence-conditioned supervision. For dense localization, \textbf{Language-guided Visual Anchoring} uses the final judgment as a compact semantic condition to construct image-specific normal and abnormal anchors, which are contrasted with dense visual features to produce anomaly maps. This keeps language as semantic guidance while grounding pixel-level responses in visual evidence. Under a strict cross-dataset zero-shot protocol on five IAD benchmarks, \method outperforms the evaluated LVLM-based baselines on most detection, localization, and judgment metrics while remaining competitive with CLIP-based methods. Ablations further validate both mixed-trajectory supervision and final-judgment conditioning.