发表机构
Singapore University of Technology and Design; Agency for Science, Technology and Research (A*STAR); Nanyang Technological University; Chongqing University(新加坡科技设计大学; 新加坡科技研究局; 南洋理工大学; 重庆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态小语言模型现场部署的鲁棒性评估问题,开发RobustMAD基准测试,通过多种开放式查询评估模型,发现最佳模型虽有潜力但仍存差距及三种失败模式,为下一代多模态工业检测助手设计提供指导。
AI 中文摘要
多模态工业异常检测助手是下一代智能工厂的关键组成部分,可实现基于视觉与语言的交互式查询。然而,由于计算需求过高和基于云推理的隐私风险,多模态大语言模型在现场部署中不切实际。紧凑的多模态小语言模型提供了可部署的替代方案,但由于缺乏全面的鲁棒性分析和反映现实工业条件的具有挑战性的基准测试,进展受到限制。为解决这一差距,我们开发了RobustMAD,这是首个基于部署动机的基准测试,旨在通过涵盖对象理解、异常检测、无法回答的问题和视觉质量退化等多种开放式查询来全面评估模型鲁棒性。与传统假设相反,表现最佳的多模态小语言模型展现出有前景的能力,甚至超过了更大的GPT-5 Nano。然而,它们仍未达到安全关键要求,RobustMAD揭示了存在操作风险的关键鲁棒性差距。特别是出现了三种反复出现的失败模式:(i)在细粒度区分或视觉条件退化下脆弱的多模态基础;(ii)响应不够全面;(iii)在无法回答或不适定查询上的逻辑基础薄弱,导致产生幻觉输出。基于这些见解,我们为利用其有前景能力的下一代多模态工业检测助手的设计提供了可操作的指导。代码可在该https URL获取。
英文摘要
Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying. However, multimodal large language models remain impractical for on-site deployment due to prohibitive computational demands and privacy risks from cloud-based inference. Compact multimodal small language models (MSLMs) offer a deployable alternative, yet progress is constrained by the lack of comprehensive robustness analyses and meaningfully challenging benchmarks that reflect real-world industrial conditions. To address this gap, we develop RobustMAD, the first deployment-motivated benchmark, designed to comprehensively evaluate model robustness through diverse open-ended queries spanning object understanding, anomaly detection, unanswerable problems, and visual quality degradations. Contrary to conventional assumptions, top-performing MSLMs exhibit promising capabilities, surprisingly outperforming even the larger GPT-5 Nano. However, they still fall short of safety-critical requirements, and RobustMAD reveals critical robustness gaps that pose operational risks. In particular, three recurring failure modes emerge: (i) fragile multimodal grounding under fine-grained distinctions or degraded visual conditions, (ii) insufficiently comprehensive responses, and (iii) weak logical grounding on unanswerable or ill-posed queries, leading to hallucinated outputs. Grounded in these insights, we provide actionable guidance for the design of next-generation multimodal industrial inspection assistants that leverage their promising competence. Code is available at https://github.com/en-research/RobustMAD.
CommentsAccepted for publication in Transactions on Machine Learning Research (TMLR), 2026 at https://openreview.net/forum?id=skrA9UYNIZ