arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05880cs.LGcs.AI

基于证据图约束的管道仪表图可靠推理:利用恢复的证据图约束视觉语言模型

Grounded and Faithful P&ID Reasoning: Constraining Vision-Language Models with Recovered Evidence Graphs

Prathamesh Gadekar, Sagar Srinivas Sakhinana, Venkataramana Runkana

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉语言模型在管道仪表图推理中易臆造连接的问题,提出恢复显式证据图并仅通过只读算子查询作答的方法,在TopoPID-VQA上将精确匹配准确率从约37%提升至约75%。

中文摘要 AI 辅助

管道仪表图(P&ID)是过程工厂的权威图谱:隔离、维护和HAZOP决策都依赖于设备之间的连接关系。视觉语言模型能够流畅地描述这些图纸,但它们常常臆造或遗漏过程连接——而一个臆造或遗漏的连接可能逆转隔离或可达性判断,因此工厂决策无法信任未经图纸线条核实的流畅回答。我们转而恢复图纸的显式图结构——其符号、符号之间的过程连接以及命名它们的标签——然后要求模型仅通过七个只读算子查询该图来作答,从而仅在引用支持其结论的查询结果时才返回拓扑断言。在TopoPID-VQA(一个包含3000个关于这些图纸的拓扑问题的新测试集)上,图接地框架(Ours)将仅图像提示下的精确匹配准确率从36.7%–41.3%提升至Qwen3-VL-4B、Qwen3-VL-8B和Gemma-4-E4B的74.3%–76.0%。这是在非完美基板上实现的:在Digitize-PID数据集上,恢复的图在精确过程连接上的F1得分为0.742,在合并符号和标签后得分为0.801。残余误差与这一差距相关——接地在恢复图正确的地方发挥作用,而在恢复图不正确的地方,感知错误仍会破坏拓扑问题。

英文摘要

Piping and Instrumentation Diagrams (P&IDs) are the authoritative maps of process plants: isolation, maintenance, and HAZOP decisions depend on what connects to what. Vision-language models describe these sheets fluently, yet they often invent or miss process connections---and an invented or missed link can reverse an isolation or reachability call, so a plant decision cannot trust a fluent answer that was never checked against the linework. We instead recover an explicit graph of the drawing---its symbols, the process connections between them, and the tags that name them---and then require the model to answer only by querying that graph through seven read-only operators, so a topology claim is returned only when it cites the query results that support it. On TopoPID-VQA, a new suite of 3000 topology questions over these sheets, Graph-Grounded Harness (Ours) raises exact match accuracy from 36.7--41.3% under image-only prompting to 74.3--76.0% for Qwen3-VL-4B, Qwen3-VL-8B, and Gemma-4-E4B. It does so on an imperfect substrate: on Digitize-PID dataset the recovered graph scores F1 0.742 on exact process connections, and 0.801 once symbols and tags are pooled in. The residual errors track that gap---grounding pays off where the recovered graph is right, and perception error still breaks topology questions where it is not.

发表机构

  • Tata Research Development and Design Center(塔塔研究开发与设计中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑