发表机构
Entropy AI Research Labs Private Limited(熵人工智能研究实验室私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对制造业视觉检测难题,提出答案条件思维链蒸馏法,利用最少标注数据让小型视觉语言模型适应新工业任务,经实验验证该方法在多任务中优于直接微调,提升了推理质量和模型性能。
AI 中文摘要
在制造业中部署基于人工智能的视觉检测很困难,因为需求经常变化、新缺陷类型出现且很少有大型标注数据集。我们提出答案条件思维链(CoT)蒸馏,以使用最少的标注数据快速使小型视觉语言模型(VLM)适应新的工业任务。前沿VLM接收训练图像及其正确标签并生成合理的视觉解释,然后通过LoRA在这些推理增强的示例上对3B参数模型进行微调。通过以正确答案为条件,确保所有训练推理都指向正确结论。我们在四个工业分类任务上进行验证,每个任务仅使用18至30个标注图像。我们的方法在所有16个种子任务组合上均优于直接微调,平均提高1.7至4.4个百分点。控制等预算实验证实改进来自推理质量而非额外训练步骤。无条件基线表明没有答案条件时错误推理会使性能下降17.8个百分点。在焊缝射线照片分类中,微调后的3B模型仅使用24个训练图像就比GPT-4.1高出了10.0个百分点。
英文摘要
Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled datasets are rarely available. We propose answer-conditioned chain-of-thought (CoT) distillation for rapidly adapting small vision-language models (VLMs) to new industrial tasks using minimal labeled data. A frontier VLM receives each training image along with its correct label and generates a justified visual explanation. A 3B-parameter model is then fine-tuned on these reasoning-augmented examples via LoRA. By conditioning on correct answers, we ensure all training reasoning is directed toward the correct conclusion, which is critical because frontier models score as low as 24.1% on our hardest task. We validate on four industrial classification tasks spanning three image modalities using only 18 to 30 labeled images per task. Across 4 seeds per task (32 training runs), our method outperforms direct fine-tuning on all 16 seed-task combinations, with mean improvements of +1.7 to +4.4 percentage points. A controlled equal-budget experiment confirms the improvement comes from reasoning quality, not additional training steps. An unconditioned baseline demonstrates that with out answer-conditioning, wrong reasoning degrades performance by 17.8 percentage points. On weld radiograph classification, the fine-tuned 3B model outperforms GPT-4.1 by 10.0pp using just 24 training images.
Comments11 pages, 5 figures, 8 tables