AI 中文总结
研究基于检索和微调的大语言模型方法在工业资产健康监测中的故障传感器诊断推理表现,通过FailureSensorIQ基准测试,发现语义搜索和混合搜索效果好,基于LLM的微调模型性能最佳,且各方法对一些干扰因素敏感。
AI 中文摘要
工业工厂运行着许多重要机器,工程师虽能用经验识别诊断机器问题,但将此推理能力转移到计算机系统仍困难。本文利用FailureSensorIQ基准研究仅检索方法和开源大语言模型(LLM)在故障传感器诊断推理中的表现。仅检索方法中测试比较了TF-IDF、BM25、语义搜索和混合搜索。基于LLM的方法评估了Qwen2.5-7B-Instruct模型的零样本、少样本提示和QLoRA微调。结果表明语义搜索和混合搜索优于纯关键词匹配技术,基于LLM的方法中微调模型性能最佳。错误分析显示随着答案选项数量增加性能下降,鲁棒性分析表明所有方法对选项洗牌、标签更改、释义和额外干扰项敏感。
英文摘要
Industrial plants run many important machines such as pumps, turbines, and compressors. Although engineers can use their experience to identify and diagnose machine problems, transferring this reasoning ability to computer systems remains difficult. This work studies how well a retrieval-only method and an open-source large language model (LLM) perform failure-sensor diagnostic reasoning using the FailureSensorIQ benchmark, a multiple-choice question-answering task introduced by IBM Research. In the retrieval-only approach, each answer option is converted into an option-level query and scored using similar correct and incorrect records from the training data. TF-IDF, BM25, semantic search, and hybrid search are tested and compared. In the LLM-based approach, the Qwen2.5-7B-Instruct model is evaluated using zero-shot prompting, few-shot prompting, and QLoRA fine-tuning. The results show that semantic search and hybrid search perform better than pure keyword-matching techniques, indicating that meaning-based similarity is more important for industrial failure-sensor reasoning. Among the LLM-based methods, the fine-tuned model achieves the best performance and substantially improves over zero-shot and few-shot prompting. Error analysis shows that performance decreases as the number of answer options increases. Robustness analysis also shows that all methods are sensitive to option shuffling, changed labels, paraphrasing, and additional distractors.
Comments6 pages, 4 figures, 6 tables