arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

网络摄像头眼动能否约束驾驶模型中的 mesa 目标?一项仪器精度分析

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

Lennox Anderson, Ahmed Boutar, Jonah Mulcrone, Tal Erez

arXiv 2608.08947首次发表:更新:

AI 中文总结

该研究探究能否用 WebGazer 眼动数据约束自动驾驶模型的 mesa 目标,经多组实验发现无统计显著效果,原因是 WebGazer 误差过大无法实现物体级注视归因。

AI 中文摘要

当前自动驾驶中的危险检测系统可能会形成 mesa 目标,即通过虚假关联而非真正的危险识别来实现高训练性能的习得内部目标。我们研究通过基于网络摄像头的眼动追踪(WebGazer,详见 https://webgazer.cs.brown.edu/)捕获的人眼注视模式,是否可作为特权信息来约束 mesa 目标的形成。我们收集了 388 个真实行车记录仪片段中与危险标注同步的 137663 帧级注视样本,随后在两种校准协议(9 点/45 次点击和 11 点/440 次点击)、两种模型架构(随机森林和因果 Transformer)以及每个实验的五个随机种子上,通过配对 t 检验对该假设进行测试。所有实验均未发现注视带来的统计显著提升(p 值分别为 0.919、0.578 和 0.667)。几何分析揭示了根本原因:WebGazer 报告的误差(根据配置约为 130-257 像素)超过了 93% 的检测到的危险物体尺寸(中位数为 36 像素),使得在该仪器精度下无法实现物体级别的注视归因。

英文摘要

Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition. We investigate whether human gaze patterns, captured via webcam-based eye tracking (WebGazer.js), can serve as privileged information to constrain mesa-objective formation. We collected 137,663 frame-level gaze samples synchronized with hazard annotations across 388 real dashcam clips, then test this hypothesis across two calibration protocols (9-point/45-click and 11-point/440-click), two model architectures (Random Forest and causal Transformer), and five random seeds per experiment with paired t-tests. No experiment yields a statistically significant improvement from gaze (p = 0.919, 0.578, and 0.667 respectively). A geometric analysis reveals the root cause: WebGazer's reported error (~130-257 px depending on configuration) exceeds 93% of detected hazard object sizes (median 36 px), rendering object-level gaze attribution physically impossible at this instrument precision.

Comments6 pages, 3 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑