InspectorGPT:用于全面工业异常检测的对比推理增强型视觉语言模型(VLM)
InspectorGPT: A Comparative Reasoning Enhanced VLM for Comprehensive Industrial Anomaly Detection
- Shanghai Jiao Tong University(上海交通大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出InspectorGPT,一种结合对比推理的VLM框架,通过CoT微调与GRPO优化,配合任务向量融合的InspectorGPT-Seg实现工业异常检测,在多维度性能及未见基准泛化上表现优异。
AI中文摘要:
工业异常检测是现代制造业的关键组成部分。大多数传统无监督方法依赖于对正常特征分布进行建模,这本质上限制了其对未知类别的泛化能力。为了提升泛化性,一些近期方法通过文本提示结合视觉语言模型(VLM)实现零样本检测,但我们观察到面向推理的后训练会导致异常判别能力崩溃,部分微调模型的表现甚至逊于其基础VLM。现有方法仅提供文本决策或粗糙边界框,缺乏像素级分割。人类检测异常的明确原则是将查询图像与无缺陷参考图像对比,我们受此启发提出以对比推理为核心的VLM框架InspectorGPT。给定正常参考图像和查询图像,InspectorGPT通过对比两者识别差异,并执行带详细推理的多项检测任务。我们通过思维链(CoT)微调及定制可验证奖励的分组相对策略优化(GRPO)来内化该能力。我们进一步推出用于像素级异常掩码的InspectorGPT-Seg,分割监督可提升异常判别能力但会削弱语义推理,联合训练无法平衡两者,因此我们分别训练两个分支并通过任务向量融合将其结合。大量实验表明,InspectorGPT在多维度性能及对未见基准的泛化性上表现优异,验证了对比推理在全面工业检测中的有效性。
英文摘要:
Industrial anomaly detection is a critical component of modern manufacturing. Most traditional unsupervised methods rely on modelling normal feature distributions, inherently limiting generalization to unknown categories. To improve generalizability, some recent methods incorporate vision-language models (VLMs) for zero-shot detection via text prompts. However, we observe that reasoning-oriented post-training can cause anomaly discrimination to collapse, with some fine-tuned models performing worse than their base VLMs. Existing methods also provide only textual decisions or coarse boxes, without pixel-level segmentation. A more explicit detection principle comes from human inspection: anomalies are identified by comparing a query image with a defect-free reference. Inspired by this, we propose InspectorGPT, a VLM framework centered on comparative reasoning. Given a normal reference and a query image, InspectorGPT compares them to identify discrepancies and perform multiple inspection tasks with detailed reasoning. We internalize this capability through Chain-of-Thought (CoT) fine-tuning and Group Relative Policy Optimization (GRPO) with tailored, verifiable rewards. We further introduce InspectorGPT-Seg for pixel-level anomaly masks. Segmentation supervision improves anomaly discrimination but weakens semantic reasoning, while joint training fails to balance them. We therefore train the two branches separately and combine them through task-vector fusion. Extensive experiments demonstrate superior multi-dimensional performance and generalization to unseen benchmarks, validating comparative reasoning for comprehensive industrial inspection.