TransGaze-Object:基于Transformer的真实驾驶中驾驶员注视目标预测框架
TransGaze-Object: Transformer Based Driver Gaze Object Prediction Framework in Real Driving
- Indian Institute of Technology Kanpur(印度理工学院坎普尔分校)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
提出基于Transformer的TransGaze-Object框架,直接从驾驶员面部和交通场景预测注视目标,构建UD-FSG数据集,准确率达60%,优于基于注视点关联的方法。
中文摘要 AI 辅助
驾驶员注视提供了驾驶员对周围交通的视觉注意和情境意识的信息。现有的驾驶员注视估计研究以注视区域或注视向量/注视点(PoG)来表示注视。然而,目标级注视信息通过识别被注视的目标(如车辆、行人或交通信号)提供了更具语义意义的视觉注意表示。在本研究中,我们提出了一种端到端的驾驶员注视目标预测框架TransGaze-Object,即基于Transformer的注视目标预测模型。所提出的框架首先提取面部特征,包括人脸和虹膜加权的眼部特征,以及交通目标的空间特征。然后使用基于Transformer的交叉注意力机制来计算相似度分数和注意力权重,以预测驾驶员的注视目标。为了训练该模型,我们提出了一个基准驾驶员注视数据集Urban Driving-Face Scene Gaze(UD-FSG),包含同步的驾驶员面部和交通场景图像、场景目标边界框以及以2D注视坐标和注视目标表示的注视标签。TransGaze-Object模型在注视目标预测上达到了60%的总体准确率,而将估计的注视点关联到交通目标的准确率为51%。误差分析显示,TransGaze-Object减少了交通目标(预测)与背景(真实)之间的混淆,误差率为11.68%,与基于PoG的注视目标关联的23.21%误差相比,相对降低了49.7%。总体而言,结果表明直接从驾驶员面部和交通场景信息预测注视目标,而不是估计中间的注视点并随后将其与交通目标关联,是有效的。
英文摘要
Driver gaze provides information regarding driver visual attention and situational awareness to the surrounding traffic. Existing driver gaze estimation studies represent gaze in terms of gaze zone or gaze vector/point-of-gaze (PoG). However, object-level gaze information provides a more semantically meaningful representation of visual attention by identifying attended objects, such as vehicles, pedestrians, or traffic signals. In this study, we propose an end-to-end driver gaze object prediction framework, TransGaze-Object, Transformer-based Gaze Object prediction model. The proposed framework first extracts facial features, including face and iris-weighted eye features, along with trafficobject spatial features. A transformer based cross-attention mechanism is then used to compute similarity scores and attention weights for predicting the drivers gaze object. To train this model, we propose a benchmark driver gaze dataset, Urban Driving-Face Scene Gaze (UD-FSG), comprising synchronized driver-face and traffic-scene images, scene objects bounding boxes, and gaze labels in terms of 2D gaze coordinate and gaze object. The TransGaze-Object model achieves an overall accuracy of 60% for gaze-object prediction, compared to 51% accuracy obtained from associating the estimated Point-of-Gaze to traffic objects. The error analysis reveals that TransGaze-Object reduces confusion between traffic objects (predicted) and the background (ground-truth), achieving an error rate of 11.68%, a 49.7% relative reduction compared with 23.21% error obtained from PoG-based gaze-object association. Overall, the results demonstrate the effectiveness of directly predicting gaze objects from driver-face and traffic-scene information, rather than estimating an intermediate Point-of-Gaze and subsequently associating it with traffic objects.