arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MTF-Net:用于行人意图预测的多模态时间特征融合网络

MTF-Net: Multi-Modal Temporal Feature Fusion Network for Pedestrian Intention Prediction

Md Mahfuzur Rahman, Pengzhan Zhou, A. F. M. Abdun Noor, Md Imam Ahasan, Md Mustafizur Rahman, Fang Qu

arXiv 2609.20178首次发表:更新:

发表机构

Chongqing University; Daffodil International University(重庆大学; 达芙妮国际大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有方法在单模态或粗粒度融合上的局限,提出MTF-Net,通过门控线性单元在循环框架中融合四种模态并利用多时间编码分支,在PIE和JAAD上以实时性能取得最高0.95和0.94的AUC,验证了原则性多模态融合的有效性。

AI 中文摘要

准确预测行人意图对于确保自动驾驶车辆与行人之间的安全、主动交互至关重要。然而,现有方法通常依赖于在单个模态内建模时间依赖性或仅在粗略语义级别融合模态的架构。为解决这些局限性,我们提出了MTF-Net,一种新颖的多模态时间特征融合网络,该网络联合建模运动学、外观和上下文线索以进行行人意图预测。MTF-Net在由门控线性单元(GLUs)增强的循环融合框架内,集成了四种互补模态——边界框动态、人体姿态关键点、局部上下文和场景级语义。这些基于GLU的模块自适应地调节跨模态信息流,实现跨时间尺度的可解释且高效的特征交互。通过三个专门的时间编码分支和一个注意力引导的融合头,所提出的模型能够在行人过街意图发生前几帧稳健地预测该意图。在PIE和JAAD基准上的广泛评估表明,MTF-Net超越了最近的基于Transformer和图模型的性能,在PIE上达到高达0.95的AUC,在JAAD上达到0.94的AUC,同时保持实时性能。结果强调,可靠的行人意图预测源于原则性的多模态融合,而非过度的架构复杂性。

英文摘要

Accurately predicting pedestrian intentions is crucial for ensuring safe and proactive interaction between autonomous vehicles and pedestrians. However, existing approaches often depend on architectures that either model temporal dependencies within individual modalities or fuse modalities only at coarse semantic levels. To address these limitations, we propose MTF-Net, a novel Multi-Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. MTF-Net integrates four complementary modalities-bounding-box dynamics, human pose keypoints, local context, and scene-level semantics within a recurrent fusion framework enhanced by gated linear units (GLUs). These GLU-based modules adaptively regulate cross-modal information flow, enabling interpretable and efficient feature interaction across temporal scales. Through three dedicated temporal encoding branches and an attention-guided fusion head, the proposed model robustly anticipates pedestrian crossing intentions several frames before they occur. Extensive evaluations on the PIE and JAAD benchmarks demonstrate that MTF-Net surpasses recent transformer- and graph-based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD, while maintaining real-time performance. The results highlight that reliable pedestrian intention prediction arises from principled multi-modal fusion rather than excessive architectural complexity.

Comments15 pages, 5 figures, accepted on International Conference on Cloud and Network Computing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑