发表机构
National Taiwan University; AMD GenAI(国立台湾大学; AMD生成式人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Anomaly-LR,一种在视觉潜在空间进行缺陷接地推理的框架,并构建首个潜在推理指令数据集IAD-LR-22K,在多个IAD基准上达到最先进性能。
AI 中文摘要
工业异常检测(IAD)正从传统的检测与定位向多模态检测系统演进,这类系统能够描述、解释并推理细粒度缺陷。尽管近期基于多模态大语言模型(MLLM)的方法通过文本推理和视觉引导提升了异常理解能力,但在细粒度检测中仍面临两个局限。首先,其视觉细化过程往往需要迭代地重新访问局部图像区域或借助额外工具。其次,由此产生的局部缺陷证据在后续推理过程中可能无法被可靠保留。为解决这些问题,我们提出了Anomaly-LR,一种缺陷接地的潜在推理框架,该框架首先形成对输入的整体理解,然后直接在视觉潜在空间中逐步细化与异常相关的表征。我们进一步构建了IAD-LR-22K,这是首个专为潜在推理设计的IAD指令数据集,包含来自4,523张工业图像的22,228个图像-问题实例,并配有全局文本推理轨迹和区域级视觉标注。大量实验表明,在多个IAD基准上,Anomaly-LR在同等规模方法中达到了最先进的性能,且无需外部参考或工具。代码和数据将在该https URL发布。
英文摘要
Industrial anomaly detection (IAD) is evolving beyond conventional detection and localization toward multimodal inspection systems that can describe, explain, and reason about fine-grained defects. Although recent multimodal large language model (MLLM)-based methods improve anomaly understanding through textual reasoning and visual guidance, they face two limitations in fine-grained inspection. First, their visual refinement often requires iteratively revisiting local image regions or augmenting with additional tools. Second, the resulting local defect evidence may not be reliably preserved throughout subsequent reasoning. To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space. We further construct IAD-LR-22K, the first IAD instruction dataset designed for latent reasoning, containing 22,228 image-question instances from 4,523 industrial images, with global textual reasoning traces and region-level visual annotations. Extensive experiments show that Anomaly-LR achieves state-of-the-art performance among comparable-scale methods across multiple IAD benchmarks, without requiring external references or tools. The code and data will be released at https://github.com/Yen666/Anomaly-LR.
Comments5 pages