arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17723cs.CV

用于模拟仪表读数的视觉-语言模型:专业化、迁移与可靠性的实证研究

Vision-Language Models for Analog Gauge Reading: An Empirical Study of Specialization, Transfer and Reliability

  • University of Central Florida(中佛罗里达大学)
  • Siemens Energy(西门子能源)

机构由 AI 辅助整理,请以论文原文为准。

Abdul Mueez, Aaditya Baranwal, Junior Chaj-Mejia, Guneet Bhatia, Jason T. Voelker, Shruti Vyas

AI总结:

本研究通过实证评估Qwen2.5-VL-7B-Instruct模型在模拟仪表读数任务中的表现,发现其经QLoRA微调后在三类数据集上误差较低,但存在迁移性能下降和高置信度错误问题,暂不适合直接部署为工厂监控流程。

AI中文摘要:

模拟仪表仍普遍存在于人工检查成本高昂或危险的工业环境中。本文解决的工程应用问题是直接读取单目标模拟仪表图像的数值,而人工智能方面的贡献是对通用视觉-语言模型(VLM)在专业化、迁移、鲁棒性和可靠性方面进行系统评估,该模型无需显式的指针分割和几何读数流程。本研究对Qwen2.5-VL-7B-Instruct模型进行评估,采用零样本提示、上下文学习(ICL)以及量化低秩适配(QLoRA)的参数高效微调,评估所用数据集包括公开合成数据集、源自视频的压力表数据集以及私有工业数据集。所有微调实验均采用固定的20轮训练协议,使用最后一轮进行分析;设置包含与不包含仪表量程的独立模型,以消除提示设置的混淆影响。主要评估指标为量程归一化平均百分比误差(MPE)。微调后的最佳MPE值在合成数据集上为2.39%,95%自举置信区间(CI)为1.43-3.90%;在压力表数据集上为2.61%,CI为1.66-3.80%;在私有工业数据集上为4.43%,CI为2.31-7.14%。留一数据集实验显示,在留出的合成数据和私有数据上存在显著的迁移性能下降,而鲁棒性测试确定高斯模糊是所测试的最强干扰因素。可靠性分析表明,高置信度下仍可能出现错误,这促使在安全关键型应用中采用弃权(不执行)和独立验证的策略。这些结果支持使用QLoRA微调的VLM进行直接单仪表读数,但尚未形成可部署的工厂监控流程。

英文摘要:

Analog gauges remain common in industrial environments where manual inspection is costly or hazardous. The engineering application addressed here is direct numerical reading of single-target analog-gauge images, while the artificial-intelligence contribution is a systematic evaluation of specialization, transfer, robustness and reliability for a general-purpose vision-language model (VLM) without an explicit pointer-segmentation and geometric-reading pipeline. The Qwen2.5-VL-7B-Instruct model is evaluated using zero-shot prompting, in-context learning (ICL) and parameter-efficient fine-tuning with Quantized Low-Rank Adaptation (QLoRA) on a public synthetic dataset, a video-derived Pressure Gauge dataset and a proprietary industrial dataset. All fine-tuning experiments use a fixed 20-epoch protocol with the final epoch used for analysis; separate models with and without supplied gauge ranges remove prompt-setting confounds. The primary metric is range-normalized mean percentage error (MPE). The best fine-tuned MPE values are 2.39% on the synthetic dataset, with a 95% bootstrap confidence interval (CI) of 1.43-3.90%; 2.61% on the Pressure Gauge dataset, with a CI of 1.66-3.80%; and 4.43% on the proprietary industrial dataset, with a CI of 2.31-7.14%. Leave-one-dataset-out experiments reveal substantial transfer degradation on held-out synthetic and proprietary data, while robustness tests identify Gaussian blur as the strongest tested corruption. Reliability analysis shows that high-confidence errors remain possible, motivating abstention and independent validation in safety-critical use. These results support QLoRA-specialized VLMs for direct single-gauge reading but not yet a deployment-ready plant-monitoring pipeline.

补充信息

↑