发表机构
Istituto Italiano di Tecnologia; Imperial College London; Italian National Institute for Insurance against Accidents at Work (INAIL)(意大利技术研究院; 帝国理工学院; 意大利国家工伤事故保险研究所(INAIL))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于单摄像头和图神经网络的列车驾驶员状态识别系统,结合面部与骨骼特征,在警觉/非警觉分类中达99%准确率,并引入多光照条件数据集。
AI 中文摘要
驾驶员疲劳对铁路安全构成重大挑战,传统的系统如死亡开关仅提供有限且基本的警觉性检查。本研究提出了一种基于视觉的监测系统,仅依赖单个前置RGB摄像头和图神经网络,将模拟列车驾驶员状态分类为警觉、非警觉以及包含模拟紧急行为的紧急类别。为了优化模型的输入表示,进行了消融研究,比较了三种特征配置:仅骨骼、仅面部以及两者结合。实验结果表明,在光照条件下,结合面部和骨骼特征的三类模型达到了最高准确率(81%),优于仅使用面部或骨骼特征的模型。此外,在光照条件下,面部和骨骼特征的组合在警觉/非警觉分类中达到了99%的准确率。我们还引入了一个受控的RGB视频数据集,包含在三种光照条件下记录的警觉、非警觉和模拟紧急行为。这些贡献代表了基于面部和上半身动态的被动非接触式列车驾驶员状态识别迈出的一步。
英文摘要
Driver fatigue poses a significant challenge to railway safety, with traditional systems like the dead-man switch offering limited and basic alertness checks. This study presents a vision-based monitoring system that relies solely on a single front-facing RGB camera and a graph neural network to classify simulated train-driver states into alert, not-alert, and an emergency class comprising acted emergency-like behaviours. To optimize input representations for the model, an ablation study was performed, comparing three feature configurations: skeletal-only, facial-only, and a combination of both. Experimental results show that combining facial and skeletal features yields the highest accuracy (81%) for the three-class model under the light condition, outperforming models that use only facial or skeletal features. Furthermore, the combination of facial and skeletal features achieves 99% accuracy in the alert/not alert classification in light condition. Additionally, we introduced a controlled RGB video dataset containing alert, not alert, and acted emergency-like behaviours recorded under three illumination conditions. These contributions represent a step toward passive and non-contact train-driver state recognition based on facial and upper-body dynamics.
Comments11 pages,5 figurees