AI 中文总结
研究针对P&IDs数字化难题,提出基于多模态大语言模型的两阶段工作流程,将其数字化重定义为设备标签提取与拓扑推理,经案例研究评估,该方法比端到端数字化更优,凸显知识引导工作流程在P&ID数字化中的潜力。
AI 中文摘要
管道和仪表图(P&IDs)编码了过程工厂的功能结构,是数字孪生和智能决策支持中关键但未充分利用的工程知识来源。然而,由于绘图标准的异构性以及现有方法对脆弱符号识别和基于规则的连接性重建的依赖,将传统P&IDs数字化仍然具有挑战性。这项工作将P&ID数字化重新定义为设备标签提取和过程拓扑推理,而非图形复制。我们提出了一种基于多模态大语言模型的两阶段工作流程,其中视觉提取和拓扑重建被视为由化学工程过程知识指导的不同推理阶段。该方法在两个复杂度不断增加的ANSI标准P&ID案例研究上进行了评估。结果表明,与端到端数字化相比,分解视觉提取和拓扑推理能产生更准确且结构一致的过程表示,突出了基于语言模型、知识引导的工作流程在可扩展和语义可靠的P&ID数字化方面的潜力。
英文摘要
Piping and instrumentation diagrams (P&IDs) encode the functional structure of process plants and are a critical yet underutilised source of engineering knowledge for digital twins and intelli-gent decision support. However, digitising legacy P&IDs remains challenging due to heterogene-ous drawing standards and the reliance of existing methods on brittle symbol recognition and rule-based connectivity reconstruction. This work reframes P&ID digitization as the extraction of equipment tags and inference of process topology, rather than graphical reproduction. We pro-pose a two-stage workflow based on multimodal large language models, in which visual extrac-tion and topology reconstruction are treated as distinct reasoning stages guided by chemical en-gineering process knowledge. The approach is evaluated on two ANSI-standard P&ID case stud-ies of increasing complexity. Results show that decomposing visual extraction and topology rea-soning yields more accurate and structurally consistent process representations than end-to-end digitization, highlighting the potential of language-model-based, knowledge-guided workflows for scalable and semantically reliable P&ID digitization.