发表机构
Louisiana State University; University of Illinois Urbana-Champaign(路易斯安那州立大学; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DrawingVQA针对建筑图纸多深度视觉文本推理设计基准测试,利用图纸与问答对,提出双分类框架评估模型,揭示模型与专家表现差距,为领域特定多模态推理及AI与工程工作流程整合奠定基础。
AI 中文摘要
我们引入了DrawingVQA,这是首个旨在评估多模态大语言模型在真实世界建筑图纸上的基准测试,建筑图纸是建筑、土木及其他工程实践中的核心媒介。与自然图像或示意性平面图不同,建筑图纸融合了抽象几何、符号标注、表格数据、注释和特定领域文本,构成独特复杂的视觉文本领域核心。DrawingVQA利用33张“施工发行版”图纸和92个精心策划的问答对弥合差距,涵盖三个推理深度。为评估模型能力,我们提出双分类框架。对先进多模态大语言模型的评估揭示了模型与专家表现之间的巨大差距。该基准测试为领域特定多模态推理奠定基础,推动人工智能驱动的理解与现实工程工作流程的整合。
英文摘要
We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.
CommentsCVPR 2026 Findings accepted paper