TrajFusionNet+:基于轨迹表示与场景图融合的Transformer行人过街意图预测
TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs
浏览论文内容
中文总结 AI 辅助
提出TrajFusionNet+,一种基于Transformer的三分支融合模型,结合轨迹序列、视觉与场景图表示,在PIE和JAAD数据集上实现最先进性能,并引入联合训练分别评估的新协议以提升泛化能力。
中文摘要 AI 辅助
行人过街意图任务涉及从自动驾驶车辆的角度预测行人是否可能横穿道路。我们提出了TrajFusionNet+,一种新颖的基于Transformer的行人过街意图预测模型。TrajFusionNet+将行人轨迹的序列和视觉表示与场景上下文的基于图的表示相结合,以预测行人过街意图。所提出的架构建立在我们之前的模型TrajFusionNet之上,包含三个分支:序列注意力模块(SAM),处理过去和预测行人轨迹的序列表示;视觉注意力模块(VAM),通过将观测和预测的边界框叠加到场景图像上,利用行人轨迹的视觉表示;以及图注意力模块(GAM),从分割的场景图像中提取以行人为中心的图,并捕捉行人与交通元素之间的关系依赖。TrajFusionNet+在两个最广泛使用的行人过街意图数据集PIE和JAAD上取得了改进的最先进性能。此外,我们引入了一种新的评估协议,在该协议中,模型在PIE和JAAD数据集上联合训练,但在每个数据集上分别评估。在这种设置下,TrajFusionNet+相比现有方法表现出更优的泛化能力。
英文摘要
The pedestrian crossing intention task involves predicting whether pedestrians are likely to cross the road from the point of view of an autonomous vehicle. We introduce TrajFusionNet+, a novel transformer-based model for pedestrian crossing intention prediction. TrajFusionNet+ combines sequential and visual representations of pedestrian trajectory with a graph-based representation of the scene context in order to predict pedestrian crossing intention. The proposed architecture builds upon our previous model, TrajFusionNet, and comprises three branches: a Sequence Attention Module (SAM), which processes a sequential representation of past and predicted pedestrian trajectories; a Visual Attention Module (VAM), which utilizes a visual representation of the pedestrian trajectories by overlaying observed and predicted bounding boxes onto scene images; and a Graph Attention Module (GAM), which extracts pedestrian-centric graphs from segmented scene images and captures the relational dependencies between pedestrians and traffic elements. TrajFusionNet+ achieves improved state-of-the-art performance on the two most widely used pedestrian crossing intention datasets, PIE and JAAD. Furthermore, we introduce a new evaluation protocol in which models are trained jointly on the PIE and JAAD datasets but evaluated separately on each. Under this setting, TrajFusionNet+ demonstrates superior generalization compared to existing approaches.
发表机构
- Perception, Robotics and Intelligent Machines Research Group (PRIME)(感知、机器人与智能机器研究组(PRIME))
- Department of Computer Science, Université de Moncton(蒙克顿大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。