arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需训练的人机协作异常检测:基于记忆库修正

Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction

Ayusha Abbas, Saram Abbas, Kabita Adhikari

arXiv 2608.17775首次发表:更新:

发表机构

School of Engineering, Newcastle University(纽卡斯尔大学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无需训练的人机协作异常检测框架,通过专家修正PatchCore检测器的记忆库,仅用少量黄金样本结合修正,可改进MVTec AD多数类别异常检测效果,成本低于全面审查。

AI 中文摘要

异常检测器最难部署的场景,恰恰是训练数据最稀缺的地方:一条新投产的生产线仅有少量经过验证的“黄金”样本,且工厂现场没有机器学习工程师。我们提出了一种无需训练的人机协作框架,领域专家可通过直接编辑记忆库来修正PatchCore检测器:无需重新训练、无需梯度、无需原始训练数据。对于误报修正,我们通过一个自校准新奇性门插入经审查图像的正常块,该门仅允许那些超过池正常最近邻距离中位数的块进入。从仅基于10个黄金样本构建的记忆库来看,操作员的修正可缩小与未修正的完全训练记忆库之间中位数66%的差距(平均80%,其中三类修正后达到同等水平),显著改进了MVTec AD数据集15个类别中的12个,且未损害任何类别:10个样本加上修正效果优于数百个无修正的样本。对于已训练的记忆库,改进空间较小,且集中在记忆库对正常外观采样不足的类别(经门控后:牙刷+0.10、金属螺母+0.09、拉链+0.05、螺丝+0.05),除网格类别外,其他类别均无显著损害。评估采用保留协议(每个类别20个拆分,经Holm校正的Wilcoxon检验),因为进入记忆库的修正图像会因记忆效应使朴素评估偏向AUROC 1.0。被动和主动查询在统计上无差异;匹配标签预算对照实验表明,收益来自部署时标签生成,其成本为全面审查的43%;缺陷记忆扩展方法效果不佳。反馈由真实值模拟,而现场专家试验(在小记忆库上错误标记成本最高)仍为未来工作。

英文摘要

Anomaly detectors are hardest to deploy exactly where training data is scarcest: a newly commissioned production line has a handful of verified "golden" samples and no machine-learning engineer on the factory floor. We present a training-free human-in-the-loop framework in which a domain expert corrects a PatchCore detector by direct memory bank editing: no retraining, no gradients, no original training data. A false-positive correction inserts the reviewed image's normal patches through a self-calibrating novelty gate admitting only those beyond the median pool-normal nearest-neighbour distance. From a bank built on only ten golden samples, operator corrections close a median 66% of the gap to an uncorrected fully trained bank (mean 80%, raised by three categories that overshoot parity), significantly improving 12 of 15 MVTec AD categories and harming none: ten samples plus corrections outperform hundreds of samples without them. On already-trained banks the headroom is smaller and concentrated where the bank undersamples normal appearance (gated: toothbrush +0.10, metal nut +0.09, zipper +0.05, screw +0.05), and no category except grid is significantly harmed. Evaluation uses a held-out protocol (20 splits per category, Holm-corrected Wilcoxon), because corrected images entering the bank inflate naive evaluation toward AUROC 1.0 by memorisation. Passive and active querying are statistically indistinguishable; a matched-label-budget control attributes gains to deployment-time label production at 43% of exhaustive-review cost; a defect-memory extension fails decisively. Feedback is simulated from ground truth; live expert trials, where mislabelling is costliest on small banks, remain future work.

Comments15 pages, 9 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑