AI 中文总结
本文提出RobustDefect-LLM框架,在NEU-DET数据集上评估四种CNN,选MobileNetV3-Large实现高准确率,结合Grad-CAM与置信度路由,集成AI辅助报告,验证了工业缺陷检测工作流的可行性。
AI 中文摘要
本文提出RobustDefect-LLM,这是一种集成了深度学习分类、面向操作员的视觉证据、置信度感知决策支持、受控AI辅助报告、可追溯存储及移动交互的工业表面缺陷检测框架,可应用于统一的质量控制工作流。此处的“鲁棒性感知”指在受控图像退化条件下进行显式评估及置信度感知审核路由,而非内在鲁棒性保证。本文在六类NEU-DET数据集的1799张图像上评估了四种基于迁移学习的卷积神经网络:ResNet50、EfficientNet-B0、DenseNet121和MobileNetV3-Large,采用固定的训练、验证及保留的域内测试划分。MobileNetV3-Large取得最高的测试数值准确率(99.26%)和宏F1值(0.9926),其bootstrap 95%准确率置信区间为0.9815-1.0000。精确配对McNemar检验显示其与DenseNet121无显著差异(p=1.000)。所选模型在CPU上的单次前向传播平均耗时0.060秒(16.66 FPS)。在合成退化组合作用下,轻度强度时准确率降至87.78%,强度更高时低于40%,表明其对严重图像质量退化敏感。Grad-CAM提供视觉证据,而置信度低于0.90或Top-2边际低于0.10的预测会被路由至人工审核。这一保守策略提供了12.22%的自动覆盖度,在33个符合条件的案例中达到100%的观测选择性准确率(95%置信区间:89.43%-100.00%),同时将所有观测到的分类错误都路由至审核。在标称受控条件下,生成的100份报告全部通过确定性一致性检查,平均延迟为1.66秒。结果表明该集成工作流具备可行性,同时强调需进行校准、重复评估及实际工业验证。
英文摘要
This paper presents RobustDefect-LLM, an industrial surface-defect inspection framework integrating deep-learning classification, operator-facing visual evidence, confidence-aware decision support, controlled AI-assisted reporting, traceable storage, and mobile interaction in a unified quality-control workflow. Here, robustness-aware denotes explicit evaluation under controlled image degradation and confidence-aware review routing, not an intrinsic robustness guarantee. Four transfer-learning-based convolutional neural networks, ResNet50, EfficientNet-B0, DenseNet121, and MobileNetV3-Large, were evaluated on 1,799 images from the six-class NEU-DET dataset using fixed training, validation, and held-out in-domain test partitions. MobileNetV3-Large achieved the highest numerical test accuracy (99.26%) and macro F1-score (0.9926), with a bootstrap 95% accuracy CI of 0.9815-1.0000. An exact paired McNemar test found no significant difference from DenseNet121 (p = 1.000). The selected model averaged 0.060 s per CPU forward pass (16.66 FPS). Under combined synthetic degradation, accuracy fell to 87.78% at mild intensity and below 40% at stronger intensities, revealing sensitivity to severe image-quality deterioration. Grad-CAM supplied visual evidence, while predictions with confidence below 0.90 or a top-2 margin below 0.10 were routed to HUMAN REVIEW. This conservative policy provided 12.22% automatic coverage and 100% observed selective accuracy among 33 eligible cases (95% CI: 89.43%-100.00%), while routing both observed classification errors to review. Under nominal controlled conditions, all 100 generated reports passed deterministic consistency checks, with a mean latency of 1.66 s. Results support the feasibility of the integrated workflow while emphasizing the need for calibration, repeated evaluation, and real-world industrial validation.
Comments27 pages, 11 figures, 16 tables, preprint manuscript, the source code is publicly available