arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SafeSceneReason:连接工业危害与事故知识的多模态推理基准

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang, Ziliang Wang, Faqiang Qian, Yunli Yang, Weiyang Shi, Qibing Ren

arXiv 2608.09230首次发表:更新:

AI 中文总结

SafeSceneReason是关联工业危害与事故知识的多模态推理基准,含两类问答对,评估发现现有视觉-语言模型在工业安全推理上存显著弱点。

AI 中文摘要

工业安全理解不仅需要检测工人、设备和个人防护装备,模型还必须评估合规性、识别危险交互、解释潜在事故机制并推荐预防措施。现有安全数据集主要聚焦视觉感知或孤立的违规识别,对基于证据的推理提供的监督有限。我们推出SafeSceneReason,这是一个多模态工业安全推理基准及配套训练语料库,将工作场所场景与职业事故调查知识关联起来。SafeSceneReason结合了两条互补的数据构建流程:以场景为中心的流程将带注释的工作场所图像转换为可执行的安全场景图,并通过对对象、关系和安全规则的程序执行生成确定性答案;以报告为中心的流程从事故报告中提取图表和上下文证据,利用证据图、明确的信息边界、多步推理路径和迭代验证构建多模态问题。最终资源包含110581个经过验证的以场景为中心的问答对和13114个经过优化的以报告为中心的问答对,涵盖感知、空间与定量推理、合规性评估、证据综合、因果分析以及缓解导向的决策制定。对代表性专有和开源视觉-语言模型的评估显示,在比较、技术及多证据推理方面存在显著性能差异和持续的弱点,表明强大的通用视觉理解尚不保证可靠的工业安全推理。

英文摘要

Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce SafeSceneReason, a multimodal industrial-safety reasoning benchmark and companion training corpus that connects workplace scenes with knowledge from occupational accident investigations. SafeSceneReason combines two complementary data-construction pipelines. The scene-centric pipeline converts annotated workplace images into executable safety scene graphs and generates deterministic answers through program execution over objects, relations, and safety rules. The report-centric pipeline extracts figures and contextual evidence from accident reports and constructs multimodal questions using evidence graphs, explicit information boundaries, multi-step reasoning paths, and iterative verification. The resulting resource contains 110,581 verified scene-centric question--answer pairs and 13,114 refined report-centric question--answer pairs, covering perception, spatial and quantitative reasoning, compliance assessment, evidence synthesis, causal analysis, and mitigation-oriented decision making. Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑