AI 中文总结
本研究通过五个工业诊断数据集发现,LLM作为集成层时,外部专家信息虽能提升其性能,但集成输出仍不及独立专家,未实现潜在互补性,强调应分别评估来源质量与集成质量。
AI 中文摘要
大型语言模型越来越多地被用作专业工具之上的集成层,但更强的组件并不必然产生更强的组合系统。在五个诊断数据集(轴承振动、过程监控、半导体设备)上,我们研究了LLM能否可靠地使用外部诊断信息;配对重复调用将建议效应与输出不稳定性分离。在所有五个数据集中,冲突的外部信息推翻了最初正确的LLM判断。在四个具有直接集成比较的数据集中,没有一个显示出隐式LLM集成相对于更强的独立来源具有一致优势。在评估前协议已固定的Tennessee Eastman确认集上,无辅助准确率为64.67%,隐式LLM-专家集成为77.43%,而专家单独为83.33%。专家信息将LLM提升了12.8个百分点(95%区间9.7至15.9),但集成输出仍比专家低5.9个百分点(95%区间-12.0至-0.7)。一个双源选择器预言机达到92.76%,表明集成输出未完全实现的互补性。集成输出错过了295个专家修正中的140个(47.5%),但丢失了99个最初正确的LLM判断中的15个(15.2%)。在提示和专家敏感性分析下,该缺陷仍然存在。在两种证据呈现下解决的CWRU案例中,任务对齐的物理证据在七个模型中的六个中产生了对错误建议易感性的较低估计(五个区间排除零);更高的推理努力在五个模型中没有可靠地减少,而一个单独的四个模型TEP分析没有明确证据表明它解决了集成问题。来源质量和集成质量应分别评估:集成层应与其更强的独立组件进行比较,而不仅仅是与无辅助的LLM比较。
英文摘要
Large language models are increasingly used as integration layers above specialized tools, but a stronger component does not necessarily produce a stronger combined system. Across five diagnostic datasets (bearing vibration, process monitoring, semiconductor equipment), we study whether an LLM can reliably use external diagnostic information; paired repeat calls separate advice effects from output instability. In all five, conflicting external information overturned initially correct LLM judgments. Among the four datasets with direct integration comparisons, none showed a consistent advantage for implicit LLM integration over the stronger standalone source. On a Tennessee Eastman confirmation set whose protocol was fixed before evaluation, unaided accuracy was 64.67%, implicit LLM-specialist integration 77.43%, and the specialist alone 83.33%. Specialist information improved the LLM by 12.8 points (95% interval 9.7 to 15.9), yet the integrated output stayed 5.9 points below the specialist (95% interval -12.0 to -0.7). A two-source selector oracle reached 92.76%, indicating complementarity that the integrated output did not fully realize. The integrated output missed 140 of 295 specialist corrections (47.5%) but lost 15 of 99 initially correct LLM judgments (15.2%). The deficit remained under prompt and specialist sensitivity analyses. Among CWRU cases solved under both evidence presentations, task-aligned physical evidence yielded lower estimates of susceptibility to incorrect advice in six of seven models (five intervals excluding zero); higher reasoning effort gave no reliable reduction in five models, and a separate four-model TEP analysis gave no clear evidence that it resolves the integration problem. Source quality and integration quality should be evaluated separately: an integration layer should be compared with its stronger standalone component, not only with the unaided LLM.
Comments35 pages, 5 figures + 1 appendix figure