AISA:用于公路施工持续改进的AI安全助手框架
AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction
浏览论文内容
中文总结 AI 辅助
本研究提出AISA框架,基于LLM和文本嵌入模型实现公路施工事故分类、质量评分及相关数据检索,为未来智能体应用提供基础,在OIICS分类和事故检索等任务上取得优于随机水平的效果。
中文摘要 AI 辅助
作业安全分析(JSA)与任务前规划可从过往事故记录中获益,但历史事故数据常以非结构化叙事形式存储,难以在规划时查询。本文提出一种以大语言模型(LLM)为核心的公路施工安全报告与规划框架,作为未来智能体应用的基础,优先采用确定性本地推理。该框架的首要目标是对现有及未来报告的事故叙事进行分类与质量评分;其次是评估相关历史事故、关联图像及可信行业文档的检索效果,以纳入日常安全计划。研究人员训练了神经探针,使其能按职业伤害与疾病分类系统(OIICS)的4个多分类字段和2个二分类字段对事故进行分类,并生成整体质量评分,在包含15000余条叙事的测试集及100条作者标注的保留集上进行评估,与多数投票LLM集成模型对比。事故、参考图像及行业文档的检索效果通过标准信息检索指标在不同嵌入模型间进行基准测试。OIICS分类在保留集上达到75%的准确率,但2个二分类标记表现退化;质量评分在一个数据库上有意义,但在保留集的分布外死亡案例中出现偏差;事故检索的相关事故恢复率远高于随机水平,在词汇差异大的施工作业上表现最佳;在文档问答任务中,开源权重解码器嵌入模型优于专有模型。总体而言,本研究提供了一种基于本地推理和文本嵌入模型的新框架,用于未来智能体应用,重点在于将外部数据与JSA报告对接。
英文摘要
Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel framework centered on large language models (LLMs) for highway construction safety reporting and planning is proposed as a foundation for future agentic applications, prioritizing deterministic, local inferencing. The first aim is to enable classification and quality scoring of incident narratives for existing and future reporting purposes. The second is to evaluate retrieval of relevant historical accidents, related imagery, and trusted industry documents for incorporation into daily safety plans. Neural probes were trained to classify incidents along four multiclass and two binary Occupational Injury and Illness Classification System (OIICS) fields and to derive an overall quality score, evaluated on a test set of over 15,000 narratives and a held-out set of 100 author-labeled records, benchmarked against a majority-vote LLM ensemble. The retrieval of historical accidents, reference imagery, and industry documents was benchmarked across embedding models using standard information retrieval metrics. OIICS classification reached 75% held-out accuracy, though the two binary flags were degenerate. The quality score, while meaningful on one database, was distorted on out-of-distribution fatalities in the held-out dataset. Accident retrieval recovered relevant incidents far above chance, performing best on lexically distinct construction activities. On document question answering, an open-weight decoder embedding model surpassed proprietary models. Overall, this work provides a new framework rooted in local inferencing and text embedding models for future agentic applications, with emphasis on bridging external data to JSA reports.
发表机构
- University of Pittsburgh(匹兹堡大学)
机构由 AI 辅助整理,请以论文原文为准。