arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于确定性几何与受约束智能体视觉-语言优化的结构平面图到模型转换

Training-Free Agentic Computer Vision for Structural Component Detection in 2D Structural Framing Plans

Mohammad Talebi-Kalaleh, Qipei Mei

arXiv 2608.17237首次发表:更新:

发表机构

University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出首个无需特定任务检测器训练的结构平面图转模型框架,经100张图基准评估,各构件指标优异,可高效修正绘图错误。

AI 中文摘要

将结构框架平面图转换为可编辑的有限元模型草稿仍十分费力,且易出现转录错误。现有的建筑构件绘图理解系统依赖针对特定任务训练的神经检测器,而结构工程中的语言模型智能体仅处理文本或模型数据,而非绘图本身。本文提出了据作者所知首个将智能体视觉-语言层应用于结构构件检测及从框架平面图PDF生成模型草稿的框架,无需针对特定任务的检测器训练或微调。确定性阶段提取图元、通过尺寸比例共识估算比例、利用绘图语法识别五类实体类别并组装可编辑布局;智能体阶段则基于确定性候选、特定操作准入测试、变更层级审查及失败关闭事务提出类型化修正。评估使用作者生成的100张平面图基准:其中开发集(用于所有规则修订)和规则冻结后生成的种子不相交保留集(仅评估一次)。所有报告分数为完整框架在保留集上的端到端结果。每张图的比例估算误差均在生成器参考值的0.1%以内;柱的召回率与精确率为0.922/0.997,梁为0.886/0.990,墙为1.000/1.000,支撑为1.000/1.000,洞口为1.000/0.964。对照研究在3张开发图上重复2次损坏各3次:校准通过全部9次试验;构件修复在9次试验中的5次满足所有严格最终状态谓词;受约束审查在明确边界内修正了遗漏的框架和错误标记。保留集与开发集共享生成器,因此研究排除了独立绘制的平面图、光栅评估、分析连通性及求解器验证。

英文摘要

Converting structural framing plans into editable finite-element model drafts is labor-intensive and susceptible to transcription errors. Existing building-component recognition systems generally depend on task-specific neural detectors, whereas language-model agents in structural engineering typically operate on text or model data rather than on drawings. To the authors' knowledge, this work is the first to apply an agentic vision-language layer to structural-component detection and model drafting from framing-plan PDFs without task-specific detector training or fine-tuning. A deterministic stage extracts geometric primitives, estimates scale by dimension-ratio consensus, recognizes five entity classes using an explicit drafting grammar, and assembles an editable layout. The agentic stage constrains typed corrections through deterministic candidates, operation-specific admission tests, change-level review, and fail-closed transactions. Evaluation used an author-generated benchmark of 100 plans, divided equally between a development half used for all rule revisions and a seed-disjoint held-out half generated after the rules were frozen and evaluated once. All scores are end-to-end results for the complete framework on the held-out half. Scale estimates were within 0.1% of the generator reference for every drawing. Recall and precision were 0.922/0.997 for columns, 0.886/0.990 for beams, 1.000/1.000 for walls, 1.000/1.000 for braces, and 1.000/0.964 for openings. A controlled study repeated two corruptions three times on three development drawings. Calibration passed all nine trials, whereas member repair satisfied every strict end-state criterion in five of nine trials. Because both benchmark halves share a generator, the evaluation does not address independently drafted plans, raster input, analytical connectivity, or solver validation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑