发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对工业异常检测问题,提出ADOPD参考特权在线策略蒸馏框架,通过参考感知教师监督仅查询的学生模型,在MMAD基准零样本推理下获77.31%平均准确率,优于Qwen3-VL-4B主干及其一样本设置。
AI 中文摘要
工业异常检测(IAD)需要识别正常视觉模式的细粒度偏差。多模态大语言模型(MLLM)可在推理时通过将查询图像与参考图像对比来提升识别准确率,但这种优势依赖额外的检索与处理。本研究探究参考对比的优势是否可内化为模型参数:训练时可访问参考图像,因此可构建参考感知的教师模型来仅监督查询的学生模型;但教师模型可能更偏向基于查询线索或语言先验的合理响应,而非有效视觉信息。为此,本文提出ADOPD,即参考特权在线策略蒸馏框架:教师模型在匹配参考与不匹配参考下评估学生模型生成的输出序列,匹配参考下的教师到学生的对数比率定义了 token 级学习方向,明确学生应学习的内容;两种参考视图间的似然差距估计参考特定支持度,校准序列级权重。实验结果显示,ADOPD在MMAD基准的零样本推理下达到77.31%的平均准确率,较Qwen3-VL-4B主干模型提升6.14个百分点,且优于其一样本设置2.64个百分点;实验表明ADOPD从参考对比中学习到细粒度异常检测策略,该项目将在指定网址开放。
英文摘要
Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) can improve recognition accuracy by comparing query images with references at inference time, but these benefits rely on additional retrieval and processing. We investigate whether the benefits of reference comparison can instead be internalized in the model parameters. Access to references during training allows a reference-aware teacher to supervise a query-only student. However, the teacher may favor plausible responses based on query cues or language priors rather than valid visual information. We propose ADOPD, a reference-privileged on-policy distillation framework. The teacher evaluates student-generated rollouts under matched and mismatched references. The matched-reference teacher-to-student log-ratio defines the token-level learning direction, specifying what the student should learn. The likelihood gap between the two reference views estimates reference-specific support and calibrates the sequence-level weight. ADOPD achieves 77.31% average accuracy on the MMAD benchmark under zero-shot inference, improving the Qwen3-VL-4B backbone by 6.14 points and outperforming its one-shot setting by 2.64 points. Experiments show that ADOPD learns a fine-grained anomaly inspection strategy from reference comparison. The project will be available at https://github.com/withTai/ADOPD.