PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
PlanSightRAG:面向民用标准图纸的自动化问答与合规检查的视觉优先多模态检索增强生成模型
专题命中 视觉问答 :grounding(summary_cn,abstract);VLM(abstract,abstract_cn);分类 cs.CV
AI总结 针对民用标准图纸的OCR自动化方法会丢失关键几何信息,本文提出视觉优先多模态RAG框架PlanSightRAG,整合多向量检索与智能体架构,构建基准数据集并验证其在检索、合规检查任务上的性能,还实现了自主视觉规则 grounding。
Comments 32 pages, 9 figures, 25 tables. Preprint submitted to Automation in Construction