arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PlanSightRAG:面向民用标准图纸的自动化问答与合规检查的视觉优先多模态检索增强生成模型

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty, Shivanand Venkanna Sheshappanavar

arXiv 2608.26091首次发表:更新:

发表机构

University of Wyoming(怀俄明大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对民用标准图纸的OCR自动化方法会丢失关键几何信息,本文提出视觉优先多模态RAG框架PlanSightRAG,整合多向量检索与智能体架构,构建基准数据集并验证其在检索、合规检查任务上的性能,还实现了自主视觉规则 grounding。

AI 中文摘要

民用基础设施合规检查长期依赖工程师手动读取传统二维图纸,但基于光学字符识别(OCR)的自动化方法会剥离解读这些图纸所需的几何形状与布局信息。本文提出一种名为PlanSightRAG的视觉优先多模态检索增强生成(RAG)框架,该框架直接对图纸图像进行索引与推理,整合了ColNomic-3B多向量检索模块、智能体式Planner-Retriever-Auditor-Synthesizer架构以及MaxSim热图作为证据追踪。我们基于五个州交通部(DOT)的标准图纸(共1898页)构建了包含4056对样本的基准数据集。PlanSightRAG在零样本检索任务上达到91.47%的Recall@5,在保留的密歇根州交通部语料库上达到91.40%。在合成的参数化生成合规图纸上,我们的Qwen2.5-VL-72B管道仅在提供预先确定的规则阈值时达到100%的判定准确率,而未使用视觉语言模型(VLM)的OCR基准已达到76.4%的准确率,这一结果构成了可控的上限。最后,我们展示了自主视觉规则 grounding能力,即无需任何人工提供的规则,直接从规范语料库中提取数值限值。

英文摘要

Civil infrastructure compliance checking has long relied on engineers manually reading legacy 2D plans; however, OCR-based automation strips away the geometry and layout essential for interpreting these plans. We present a Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG. It indexes and reasons directly over plan imagery, integrates a ColNomic-3B multi-vector retrieval, an agentic Planner-Retriever-Auditor-Synthesizer, and MaxSim heatmaps as an evidence trail. We introduce a 4,056-pair benchmark from five state Departments of Transportation (DOT) standard plans (1,898 pages). PlanSightRAG achieves 91.47% Recall@5 on zero-shot retrieval, while on a held-out Michigan DOT corpus, it achieves 91.40%. On synthetic, parametrically-generated compliance drawings, our Qwen2.5-VL-72B pipeline reaches 100% verdict accuracy only when supplied a pre-resolved rule threshold, a controlled ceiling that a non-VLM OCR baseline already reaches at 76.4%. Finally, we demonstrate autonomous visual rule-grounding by extracting numeric limits directly from a specification corpus without any human-supplied rules.

Comments32 pages, 9 figures, 25 tables. Preprint submitted to Automation in Construction

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑