发表机构
Western University; National Research Council Canada(西安大略大学; 加拿大国家研究委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对车内注视跟踪丢失问题,提出因果上下文门控预测器(CCGF),融合丢失前注视历史与DINOv3场景特征,实现因果注视恢复,在自然驾驶数据上误差降低33%。
AI 中文摘要
安装在仪表板上的注视跟踪器在驾驶员进行大幅度头部旋转(包括肩部检查、后视镜观察和路口扫视)时,往往会丢失对驾驶员眼睛的视线。这些操作发生在驾驶员视觉注意力信息最为有用的时候。离线间隙填充方法可以利用两侧的观测数据重建缺失区间,但在线驾驶员监控系统不能依赖尚未发生的测量数据。因此,我们提出了因果注视恢复:在时间t时,在没有目标跟踪器在t或之后时刻的注视数据的情况下,预测不可用的注视。我们引入了因果上下文门控预测器(CCGF),它编码了丢失前60帧的注视和头部姿态历史,并将其与DINOv3场景特征相结合。一个学习到的可靠性门控控制着随着丢失进展,历史和场景表示的贡献。我们评估了两种场景条件:实时(Live),其中场景表示在跟踪器丢失期间继续更新;冻结(Frozen),其中在缺失区间内使用丢失前的最终表示。我们在来自十名驾驶员10.5小时自然驾驶数据中,对2047个符合条件的自然发生的GazeSense head_lost事件评估了CCGF。在所有记录中,head_lost占GazeSense记录时间的8.5%。来自头戴式Neon跟踪器的同步注视坐标提供监督和评估目标,但从未用作模型输入。在留一驾驶员评估下,CCGF在实时场景更新下实现了每驾驶员平均中位误差175.7像素(10.5度),相对于仅历史因果预测减少了33%。使用冻结场景输入时,误差增加到210.8像素(12.9度),表明在丢失期间获取的场景观测提供了有用的预测信息。我们将发布数据集、评估协议和因果基线。
英文摘要
Dashboard-mounted gaze trackers often lose sight of the driver's eyes during large head rotations, including shoulder checks, mirror glances, and intersection scanning. These maneuvers occur when information about the driver's visual attention is most useful. Offline gap-filling methods may reconstruct a missing interval using observations from both sides, but an online driver-monitoring system cannot rely on measurements that have not yet occurred. We therefore formulate causal gaze recovery: forecasting unavailable gaze at time t without target-tracker gaze at t or later. We introduce the Causal Context-Gated Forecaster (CCGF), which encodes a 60-frame pre-dropout history of gaze and head pose and combines it with DINOv3 scene features. A learned reliability gate controls the contribution of the history and scene representations as the dropout progresses. We evaluate two scene conditions: Live, in which the scene representation continues to update during tracker loss, and Frozen, in which the final pre-dropout representation is used throughout the missing interval. We evaluate CCGF on 2,047 eligible, naturally occurring GazeSense head\_lost events drawn from 10.5 h of naturalistic driving by ten drivers. Across all recordings, head\_lost accounts for 8.5 percent of GazeSense recording time. Synchronized gaze coordinates from a head-mounted Neon tracker provide supervision and evaluation targets but are never used as model inputs. Under leave-one-driver-out evaluation, CCGF achieves a mean per-driver median error of 175.7 px (10.5 deg) with Live scene updates, a 33 percent reduction relative to history-only causal forecasting. With Frozen scene input, the error increases to 210.8 px (12.9 deg), indicating that scene observations acquired during the dropout provide useful predictive information. We will release the dataset, evaluation protocol, and causal baselines.
Comments8 pages, 5 figures