arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于稀疏注意力的自适应跨模态融合用于行人过街意图预测

Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

Md Mahfuzur Rahman, Pengzhan Zhou, A F M Abdun Noor, Md Imam Ahasan, Kah Ong Michael Goh, S. M. Hasan Mahmud, Md Mustafizur Rahman, Kaixin Gao

arXiv 2607.12293首次发表:更新:

发表机构

Faculty of Information Science & Technology, Multimedia University(信息科学与技术学院,马来西亚多媒体大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对自动驾驶中行人过街意图预测问题,提出ADAPT多模态框架,通过五个模块联合建模视觉和运动信息,实验证明该方法优于现有技术,在准确率和实时性上取得平衡,为智能交通等应用提供有效方案。

AI 中文摘要

预测行人过街意图对自动驾驶至关重要,但现有方法常依赖单模态输入或密集多模态融合策略,无法充分捕捉互补视觉和运动信息且引入冗余跨模态交互。我们提出ADAPT多模态框架,通过五个专门模块联合建模局部和全局视觉上下文以及时间运动动态,用于准确预测行人过街意图。实验表明ADAPT优于现有方法,计算复杂度低,在JAAD和PIE数据集上取得良好结果,且推理速度快,为智能交通和自动驾驶应用在预测准确性和实时部署效率间提供了有效平衡。

英文摘要

Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or dense multimodal fusion strategies that inadequately capture complementary visual and kinematic information while introducing redundant inter-modal interactions. We propose ADAPT (Adaptive Domain-Aware Pedestrian Crossing Transformer), a multimodal framework that jointly models local and global visual context together with temporal motion dynamics for accurate pedestrian crossing intention prediction. ADAPT processes four spatially aligned visual modalities, including RGB images, local depth maps, global semantic maps, and global depth maps, together with ego-vehicle speed, pedestrian bounding boxes, and skeleton pose information through five specialized modules: a weight-shared Swin Transformer V2 backbone for visual feature extraction, a Cross-Modality Guided Attention module for hierarchical visual fusion, a Mamba-based Motion Feature Encoding module for efficient temporal modeling, a Sparse Cross-Modal Attention module that selectively preserves the most informative inter-modal interactions, and a Vision Transformer-based Temporal Feature Fusion module for sequence-level prediction. Extensive experiments on the JAAD and PIE benchmark datasets demonstrate that ADAPT consistently outperforms existing state-of-the-art methods while maintaining low computational complexity. On JAAD, the proposed method achieves an AUC of 0.73 on JAADbeh and 0.85 on JAADall, while on PIE it achieves an accuracy of 0.92 and an AUC of 0.90. Furthermore, ADAPT performs inference in only 17.23 ms per sample, offering an effective balance between predictive accuracy and real-time deployment efficiency for intelligent transportation and autonomous driving applications.

Comments17 pages, 5 figures, 4 tables. Under review at PeerJ Computer Science

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑