arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15032cs.CLcs.AIcs.CV

Handoff-H1:一种用于从建筑蓝图中提取材料工程量的协同视觉智能体系统

Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints

Bruno Chicelli, Henrique Alves, Rodrigo Anselmo, Joshua Weinberg, Felipe Lemos, Jan Baryla

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出Handoff-H1视觉智能体系统,结合专用计算机视觉模型、工具智能体与建筑知识库,在建筑蓝图工程量提取基准上,性能优于前沿模型和人类专业估算师。

中文摘要 AI 辅助

将一组建筑蓝图转换为完整的材料工程量提取结果,需要跨图纸的视觉感知、尺寸与多跳推理,以及图纸未明确说明的建筑规范依据。我们提出Handoff-H1,这是一个包含三层架构的工程量提取系统:用于提取基本元素的专用计算机视觉模型;配备图像操作和内部视觉任务工具(包括计算机视觉模型支持的计数、检测和图纸分解)的工具使用智能体;以及基于精选建筑知识库的持久化、分层结构化项目基础。我们在建筑蓝图工程量提取基准(Construction Blueprint Takeoff Benchmark)上进行评估:10套真实住宅蓝图搭配经共识验证的专家工程量提取结果,共2009个已验证条目,评分限定为影响造价估算的1348个一级材料,由大语言模型法官按各专业在材料覆盖率、25%精度下的数量准确率(P@.25)评分,并加权合成综合得分。在与原始PDF相同的评分规则下,7个前沿开源权重模型的综合得分在35至61之间;独立专业估算师与同一协调后的黄金标准对比,得分为77.6%(65.5%覆盖率,87.9% P@.25)。Handoff-H1从原始PDF端到端运行,得分为81.6%(86.1%覆盖率,78.8% P@.25):比最强的前沿智能体高出约20分,且通过达到接近人类的数量精度和人类未实现的覆盖率,超过了独立专业估算师。评估工具对Open Harbor框架公开;蓝图集和真值可应要求供研究使用。

英文摘要

Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the drawings never state. We present Handoff-H1, a takeoff system built from three layers: purpose-built computer-vision models that extract primitives; tool-using agents equipped with image operations and in-house visual-task tools, including CV-model-backed counting, detection and plan decomposition; and a persistent, hierarchically structured project foundation, grounded in a curated construction knowledge base. We evaluate on the Construction Blueprint Takeoff Benchmark: 10 real residential blueprint sets paired with consensus-validated expert takeoffs - 2,009 verified line items, restricted for scoring to the 1,348 primary-tier materials that drive an estimate - scored per trade by an LLM judge on material coverage and quantity Precision@25% (P@.25) and combined into a weighted composite. Under identical scoring from the raw PDF, seven frontier and open-weight models span composites of 35-61, and independent professional estimators - scored against the same reconciled gold standard - post 77.6% (65.5% coverage, 87.9% P@.25). Handoff-H1, working end-to-end from the raw PDF, reaches 81.6% (86.1% coverage, 78.8% P@.25): roughly 20 points above the strongest frontier agent, and above the independent estimators by pairing near-human quantity precision with coverage they do not reach. The evaluation harness is public for the open harbor framework; the blueprint sets and ground truth are available upon request for research use.

发表机构

  • Handoff AI Research(汉德夫人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑