arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BlueprintAgent:面向扫描结构蓝图仿真就绪生成的约束触发式定向重访

BlueprintAgent: Constraint-Triggered Targeted Revisits for Simulation-Ready Generation from Scanned Structural Blueprints

Zhouyuan Xu, Chen Yang, Linhao Wang, Jiansheng Fan, Chen Wang

arXiv 2609.07362首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出BlueprintAgent,一种约束触发式多模态智能体,通过将工程约束实现为验证器并触发MLLM局部定向重访,从扫描RC结构蓝图提取仿真就绪框架,在300张真实图纸上达到0.994的宏平均梁F1,显著优于基线和固定流程。

AI 中文摘要

将在役钢筋混凝土(RC)建筑蓝图转换为仿真就绪模型——即支持确定性有限元法(FEM)导出和合格工程师审查的结构化框架表示——是安全评估和抗震加固的基础,但该过程目前仍依赖人工。直接提示多模态大语言模型(MLLM)处理扫描图纸并不可靠:输出常常违反梁-柱支撑、跨数或三维连续性等工程约束。我们提出BlueprintAgent(BPA),一种用于从扫描蓝图提取仿真就绪框架的约束触发式多模态智能体。BPA将MLLM作为主要阅读器和决策者,由OCR和计算机视觉提供局部证据。其核心机制将工程约束实现为可调用的验证器,这些验证器的实体级冲突报告会触发MLLM对局部区域进行定向重访——这是一种推理时控制,不同于固定流程和自由形式的自我反思。我们在来自20个匿名RC框架项目的300张真实扫描蓝图图纸上评估了BPA,并与五个基线和六个消融进行了比较。BPA的宏平均梁F1分数达到0.994,而单MLLM零样本为0.301,固定流程为0.820;移除MLLM主导的轴线判定会使复杂多图纸项目的梁和柱F1分数大幅下降。对于密集技术图纸,工程约束最好作为实体级定向重访的触发器部署,而非作为事后输出过滤器。

英文摘要

Converting in-service reinforced-concrete (RC) building blueprints into simulation-ready models---structured frame representations that support deterministic FEM export and qualified-engineer review---underpins safety assessment and seismic retrofit, but the process remains manual. Direct prompting of a multimodal large language model (MLLM) over a scanned sheet is unreliable: outputs often violate engineering constraints on beam--column support, span count, or 3D continuity. We present BlueprintAgent (BPA), a constraint-triggered multimodal agent for simulation-ready frame extraction from scanned blueprints. BPA treats the MLLM as the primary reader and decision maker, with OCR and computer vision supplying localized evidence. Its central mechanism realizes engineering constraints as callable validators whose entity-level conflict reports trigger targeted MLLM revisits over the local region---an inference-time control distinct from fixed pipelines and free-form self-reflection. We evaluate BPA on 300 real scanned blueprint sheets from 20 anonymized RC frame projects, against five baselines and six ablations. BPA reaches a macro-averaged Beam F1 of 0.994, against 0.301 for single-MLLM zero-shot and 0.820 for a fixed pipeline; removing MLLM-led axis adjudication collapses Beam and Column F1 on complex multi-sheet projects. For dense technical drawings, engineering constraints are best deployed as triggers for entity-level targeted revisits rather than as post-hoc output filters.

CommentsAccepted to Findings of EMNLP 2026. 14 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑