arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22081cs.CL

职业事故叙述中事故过程角色分类的跨部门泛化

Structuring occupational accident narratives for cross-sector safety analysis: Transferability of accident-process role classification

  • Université Clermont Auvergne(克莱蒙奥弗涅大学)
  • Laboratoire de Mathématiques Blaise Pascal (UMR 6620 CNRS)(布莱兹·帕斯卡数学实验室(UMR 6620 CNRS))
  • LYF SAS(LYF SAS公司)
  • Institut Universitaire de France (IUF)(法国大学研究院)

机构由 AI 辅助整理,请以论文原文为准。

Aho Yapi, Pierre Latouche, Arnaud Guillin, Yan Bailly

中文总结 AI 辅助

本研究评估法语职业事故叙述中事故过程角色分类的跨部门泛化,通过任务特定适应策略在建筑部门训练后,在冶金等未见语料上达到85.6%-85.8%的平衡准确率,支持可迁移辅助编码系统开发。

中文摘要 AI 辅助

职业事故叙述包含关于工作情况、不利条件、事故事件及其后果的宝贵信息。自动结构化这些叙述可以促进大规模事故分析,并支持职业风险预防。然而,描述事故的术语和写作风格在不同部门和组织之间差异显著,这引发了关于自动化编码系统在其训练领域之外泛化能力的问题。在本文中,我们评估了法语职业事故叙述中事故过程角色分类的跨部门泛化能力。我们构建了一个专家标注语料库,其中事实单元被分类为四个角色:工作情况(A0)、明确报告的不利条件(A1)、事故事件或偏差(B)以及报告的后果(C)。角色分类器仅基于从6,040个建筑部门叙述中提取的42,244个事实单元进行开发和选择,然后在不进行重新训练或目标领域调优的情况下,在冶金和化学-塑料部门的未见语料库以及独立收集的公司语料库上进行评估。我们比较了冻结的预训练表示与任务特定微调和监督表示学习策略。结果表明,任务特定适应在跨领域迁移中始终优于冻结表示。在重复训练运行中,三种领先的任务适应策略在三个目标语料库上的平均平衡准确率介于85.6%和85.8%之间。这些发现支持开发可迁移的辅助编码系统,该系统能够一致地结构化异构的职业事故叙述,以供专家审查和跨部门预防分析。

英文摘要

Introduction: Occupational accident narratives describe work situations, unfavourable conditions, accident events, and consequences, but differences in terminology and reporting practices hinder systematic analysis across sectors and organisations. This study examined whether a model developed in one occupational sector could identify the same accident-process information in unseen sectors and reporting environments. Method: French accident narratives were segmented into factual units and expert-annotated as work situation, explicitly reported unfavourable condition, accident event or deviation, or reported consequence. Models were developed on 42,244 factual units from 6,040 construction-sector narratives and evaluated without retraining on metallurgy, chemistry-plastics, and an independently collected company corpus. We compared a TF-IDF-based lexical model, frozen pretrained text representations, and task-adapted pretrained models. Results: Average balanced accuracy was 75.0% for TF-IDF, 76.9% for frozen pretrained representations, and about 85.7% after task adaptation. Across repeated runs, leading adapted approaches showed similar overall performance, with no method consistently outperforming the others. Performance was lower and more variable on the company corpus, where transfer also involved a different reporting environment and data source. Conclusions: Accident-process roles learned from construction narratives remained identifiable in other sectors and an independent organisational setting. Practical Applications: The framework can support assisted coding, expert review, cross-sector analysis, and prevention-oriented analysis of large accident-report collections.

↑