arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边缘端低成本两阶段织物缺陷检测

Low Cost Two-Stage Fabric Defect Detection at the Edge

Rasel Hossen, Diptajoy Mistry, Mosaddek Hossain Kamal

arXiv 2608.14727首次发表:更新:

发表机构

University of Dhaka(达卡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对低收入经济体中小工厂的低成本织物缺陷检测需求,构建含YOLOv5n的两阶段级联结构,在Jetson Nano上实现端侧部署,验证了其性能并揭示了级联提速的关键影响因素。

AI 中文摘要

低收入经济体服装行业的织物检测仍以人工为主,而商用视觉系统的价格超出了大多数中小型工厂的承受能力。由于受控生产环境下缺陷较为稀疏,一种自然的应对方案是采用级联结构:用廉价的异常检测器筛查每一帧图像,仅对可疑帧调用完整检测器。我们针对4类针织织物缺陷构建了这样的级联结构,并在NVIDIA Jetson Nano上以TensorRT FP16格式端到端部署。第一阶段是紧凑卷积自编码器,配备解码器注意力门、边缘加权重构损失,以及来自冻结YOLOv5n教师模型的特征级蒸馏;第二阶段是YOLOv5n,仅对标记帧调用。在与检测器训练集不相交的249张图像基准集(20张缺陷图像、229张非缺陷图像)上,第一阶段在以召回率为优先的阈值下标记了全部20张缺陷图像(95%置信区间0.83-1.00),假阳性率为49.3%(113/229),相比普通自编码器减少了19.3%的假阳性(p=0.011)。并行流水线达到13.45 FPS,而纯YOLO串行循环仅为9.86 FPS。我们对这1.36倍提速进行分解,核心发现是91%的提速源于JPEG解码与推理的重叠,而非级联结构,级联结构在实测转发速率下仅贡献了5.1%的推理减少(p=0.534)。进一步研究显示,此处的转发受假阳性限制而非缺陷占比限制——85%的转发帧是误报,并量化了在更严格校准下可实现29%-45%的推理减少。我们将此作为警示:在未控制数据路径的情况下测量级联提速可能存在偏差,并将该系统定位为AI辅助分诊而非自主验收。

英文摘要

Fabric inspection in the garment industries of low-income economies remains largely manual, and commercial vision systems are priced beyond most small and medium mills. Because defects are sparse under controlled production, a natural response is a cascade: screen every frame with a cheap anomaly detector and invoke a full detector only on suspicious frames. We build such a cascade for four knit-fabric defect classes and deploy it end-to-end on an NVIDIA Jetson Nano with TensorRT FP16. Stage 1 is a compact convolutional autoencoder with decoder attention gates, an edge-weighted reconstruction loss, and feature-level distillation from a frozen YOLOv5n teacher; Stage 2 is YOLOv5n, invoked only on flagged frames. On a 249-image benchmark disjoint from detector training (20 defective, 229 non-defective), Stage 1 at a recall-prioritised threshold flags all 20 defective images (95% CI 0.83-1.00) at a false-positive rate of 49.3% (113/229), reducing false positives by 19.3% relative to a plain autoencoder (p=0.011). The parallel pipeline reaches 13.45 FPS against 9.86 FPS for a sequential YOLO-only loop. Our central finding comes from decomposing that 1.36x: 91% of it is attributable to overlapping JPEG decode with inference rather than to the cascade, which contributes only a 5.1% inference reduction at the measured forwarding rate p = 0.534. We further show that forwarding here is false-positive-limited rather than prevalence-limited - 85% of forwarded frames are false alarms - and quantify the 29-45% inference reduction attainable under tighter calibration. We report this as a caution for cascade speedups measured without controlling the data path, and position the system as AI-assisted triage rather than autonomous acceptance.

Comments14 pages, 10 figures, 8 tables. Deployment study on NVIDIA Jetson Nano with TensorRT FP16. Includes a decomposition showing the measured 1.36x end-to-end speedup is dominated by data-path overlap rather than by the cascade. Dataset available on Roboflow Universe

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑