arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12310cs.CV

面向细粒度工业异常理解,仅域内训练是否足够?

Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?

  • Hunan University(湖南大学)
  • University of Aberdeen(阿伯丁大学)
  • South China Normal University(华南师范大学)
  • South China University of Technology(华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingwu Zhang, Duanyang Du, Huiling Zhu, Jiayue Dai, Yixiao Liu, Guozhi Liu, Zhihan Zhang, Zijun Long

AI总结:

针对多模态工业异常理解任务,本文发现仅域内训练无法解决单一MLLM的细粒度感知缺陷,提出SiGMA多智能体框架,在MMAD基准上性能接近人类专家且可灵活接入新模型。

AI中文摘要:

单一多模态大语言模型(MLLM)在多模态工业异常理解(MM-IAU)任务中,难以同时在检测、定位、描述和推理四个子任务上均表现出色。本文研究表明,仅进行域内训练无法缩小该性能差距:在广泛使用的MM-IAU基准MMAD上,经训练的专用模型在缺陷定位任务上的准确率最高仅为75.5%,而人类专家的准确率达92.3%,部分专用模型的异常检测准确率甚至低于未训练的基础模型。同时,不同MLLM虽各有优势,但均存在细粒度感知的共性缺陷,仅将它们组合无法消除该缺陷。为此,本文提出SiGMA,一种空间接地的多智能体框架,该框架将任务分配给异构MLLM智能体和专用视觉缺陷专家:多模态搜索器提供工业知识与正常参考样本,缺陷专家将查询与参考的对比转化为校准后的异常证据,无标签可靠性控制器则根据任务相关能力和查询级证据质量权衡各来源的权重。实验结果显示,SiGMA在MMAD上的平均准确率达85.2%,比性能最强的经训练专用模型及Gemini-2.5-Pro高4.0%,与人类专家的准确率差距在1.5%以内;即便使用最多90亿参数的三个智能体,SiGMA仍能达到84.4%的准确率,且可在不重新训练的情况下接入新的MLLM。

英文摘要:

A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU). We show that in-domain training does not close this gap. On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model. Meanwhile, different MLLMs offer complementary strengths but share this weakness in fine-grained perception, so combining them alone cannot remove it. We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert. A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality. SiGMA reaches 85.2% average accuracy on MMAD, 4.0% above the strongest trained specialist and Gemini-2.5-Pro and within 1.5% of human experts. Even with three agents of at most 9B parameters, it reaches 84.4%, and new MLLMs join without retraining.

补充信息

↑