发表机构
Oslo Metropolitan University; Simula Research Laboratory; Kristiania University of Applied Sciences; University of Oslo(奥斯陆城市大学; 西穆拉研究实验室; 克里斯蒂安尼亚应用科学大学; 奥斯陆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出两个互补的扩散模型分别生成位置和速度眼动轨迹,以低成本合成视觉搜索数据,在运动学指标上优于基线,但长距离特征有待条件模型改进。
AI 中文摘要
眼动追踪数据的采集成本高昂,需要专用硬件和受控的实验室条件,并且由于隐私限制而难以共享。我们通过两个互补的去噪扩散概率模型(DDPMs)来解决这一问题,用于从视觉搜索数据中无条件生成眼动动态。两者均使用相同的FiLM条件一维U-Net,带有自注意力机制(19.35百万参数),并在来自28名参与者的8秒滑动窗口序列上进行训练。一个模型生成原始的二维注视位置序列,而另一个生成双分量速度序列;每个模型使用特定于表示的预处理、训练设置、数据划分和评估协议。两者均在三个独立的训练种子上进行评估,聚合指标以均值±标准差报告。位置空间模型在九个运动学特征上实现了0.016±0.004的平均Jensen-Shannon(JS)散度,最高的特征级均值低于0.030,注视持续时间与真实数据的差异在2%以内,弗雷歇注视距离比统计基线和马尔可夫基线低一个数量级以上。在“合成训练-真实测试”协议下,仅合成训练实现了R²=0.66±0.02,为真实数据R²点估计的82.7%。速度空间模型在速度分量、速度、对数速度和转向角上实现了0.0065的平均JS散度,最大值为0.015±0.005。重建路径长度的准确性较低(0.21±0.02对比位置空间的0.03±0.01),尽管协议不同。总体而言,无条件扩散捕获了局部眼动运动学以及短距离的时间和方向结构,而长距离属性如扫视计数和累积路径几何仍然是未来条件模型的目标。
英文摘要
Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of privacy constraints. We address this using two complementary denoising diffusion probabilistic models (DDPMs) for unconditional generation of eye-gaze dynamics from visual-search data. Both use an identical FiLM-conditioned one-dimensional U-Net with self-attention (19.35,M parameters), trained on 8,s sliding-window sequences from 28 participants. One model generates raw two-dimensional gaze-position sequences, while the other generates two-component velocity sequences; each uses representation-specific preprocessing, training settings, data partitions, and evaluation protocols. Both are evaluated across three independent training seeds, with aggregated metrics reported as mean,$\pm$,SD. The position-space model achieves a mean Jensen-Shannon (JS) divergence of $0.016\pm0.004$ across nine kinematic features, with the highest feature-wise mean below $0.030$, fixation duration within 2% of real data, and a Fr'echet Gaze Distance more than an order of magnitude below statistical and Markovian baselines. Under a Train-on-Synthetic-Test-on-Real protocol, synthetic-only training achieves $R^2=0.66\pm0.02$, or 82.7% of the real-data $R^2$ point estimate. The velocity-space model achieves a mean JS divergence of $0.0065$ across velocity components, speed, log-speed, and turning angle, with a maximum of $0.015\pm0.005$. Reconstructed path length is less accurate ($0.21\pm0.02$ versus $0.03\pm0.01$ in position space), although the protocols differ. Overall, unconditional diffusion captures local gaze kinematics and short-range temporal and directional structure, while long-range properties such as saccade counts and cumulative path geometry remain targets for future conditioned models.
Comments14 pages, 7 figures