arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23507cs.CV

检测手机引发的行人分心:一种多模态融合Transformer

Detecting Phone-Induced Pedestrian Distraction via a Multimodal Fusion Transformer

  • Technische Universität Berlin(柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuanzhe Li, Hounian Liu, Xiaotong Chang, Yidi Huang

AI总结:

针对手机引发的行人分心问题,提出多模态融合Transformer(MFT),融合骨骼动态与视觉特征,利用跨模态和时间注意力,在287个行人实例数据集上达到95%准确率,超越六种基线6%。

AI中文摘要:

对手机的日益依赖使得手机引发的行人分心现象愈发普遍。发短信、观看视频和打电话等活动已成为交通事故的重要诱因。可靠的行人分心检测对于自动驾驶汽车至关重要,因为它能提升态势感知能力,实现及时的风险评估,从而支持安全的运动规划和车辆控制。我们提出了一种用于检测手机引发的行人分心的多模态融合Transformer(MFT)。MFT同时从人体姿态关键点中提取骨骼动态信息,并从行人图像中提取视觉外观特征,有效利用两种模态提供的互补信息。我们提出了一种跨模态注意力模块,通过多头交叉注意力捕获模态间的依赖关系,促进两种模态间互补信息的有效融合。随后,采用由Transformer编码器实现的时间注意力融合模块来捕获时间依赖关系。MFT在一个手动标注的数据集上进行了训练和评估,该数据集包含287个行人实例,共20,741张图像。大量实验表明,MFT达到了95%的总体准确率,比六种基线方法的性能高出6%。

英文摘要:

The increasing reliance on mobile phones has made phone-induced pedestrian distraction increasingly prevalent. Activities such as texting, watching videos, and making phone calls have become significant contributors to traffic accidents. Reliable detection of pedestrian distraction is essential for autonomous vehicles, as it improves situational awareness and enables timely risk assessment, thereby supporting safe motion planning and vehicle control. We propose a multimodal fusion Transformer (MFT) for detecting phone-induced pedestrian distraction. MFT jointly extracts skeletal dynamics from body pose keypoints and visual appearance features from pedestrian images, effectively leveraging the complementary information provided by the two modalities. A cross-modal attention module is proposed to capture inter-modal dependencies through multi-head cross-attention, facilitating effective fusion of complementary information across the two modalities. Then, a temporal attention fusion module, implemented with a Transformer encoder, is employed to capture temporal dependencies. MFT is trained and evaluated on a manually annotated dataset comprising 287 pedestrian instances with 20,741 images. Extensive experiments demonstrate that MFT attains an overall accuracy of 95%, exceeding the performance of six baseline approaches by 6%.

补充信息

↑