arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00068cs.CV

SafeBuild-Bench:一种具有图增强数据挖掘能力的时序鲁棒建筑安全基准

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

Yi Cui, Zilin Wang, Yijie Xu, Qianyi Cai, Huizai Yao, Shuai Jiang, Bingzhuo Zhong, Hui Xiong

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员构建了图增强数据挖掘的时序鲁棒建筑安全基准SafeBuild-Bench,开发可扩展验证的GEMS流水线,发现现有多模态大语言模型对建筑安全理解仍不足。

中文摘要 AI 辅助

建筑安全模型必须应对实际部署中的风险,例如工人站在无护栏的脚手架边缘,而非仅识别精选图像中的常见物体。然而,实际检查档案存在冗余、长尾分布的问题,且收集自不断变化的场地和数月间。我们推出SafeBuild-Bench,这是一个元数据驱动的基准,用于评估多模态大语言模型在真实时序和场地变化下的建筑安全表现。该基准从10万+工业图像-文本记录中挖掘而来,包含来自3000多张专家验证图像的3314个任务实例,涵盖多项选择题式危害识别和自由形式危害描述。为让专家验证具备可扩展性,我们开发了GEMS,这是一种图增强多模态选择流水线,结合代理模型的混淆信号与基于图的多样性,从冗余流中识别信息丰富的候选样本。在公开指令调优数据上,GEMS选定的子集在小数据预算下仍保持面向鲁棒性的性能。在SafeBuild-Bench上,当前多模态大语言模型(MLLMs)对建筑安全的理解仍远未达到可靠水平,最佳综合得分接近60。我们在该httpsURL发布了基准、评估脚本及GEMS代码库。

英文摘要

Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only recognize common objects in curated images. Yet real inspection archives are redundant, long-tailed, and collected across changing sites and months. We introduce SafeBuild-Bench, a metadata-driven benchmark for evaluating multimodal large language models on construction safety under realistic temporal and site variation. It is mined from 100K+ industrial image-text records and contains 3,314 task instances from over 3,000 expert-verified images, covering multiple-choice hazard identification and free-form hazard description. To make expert verification scalable, we develop GEMS, a graph-enhanced multimodal selection pipeline that combines a proxy-model confusion signal with graph-based diversity to identify informative candidates from redundant streams. On public instruction-tuning data, GEMS-selected subsets preserve robustness-oriented performance under small data budgets. On SafeBuild-Bench, current MLLMs remain far from reliable construction-safety understanding, with the best overall score near 60. We release the benchmark, evaluation scripts, and GEMS codebase at https://github.com/safebuild/gems.

发表机构

  • Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑