arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38519cs.CVcs.LG

GazeFlow:从人类注视行为到生成式自我中心注视预测

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

Sheng Zhao, Weikai Lin, Yuhao Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

GazeFlow利用条件流匹配将注视建模为受自上而下和自下而上信息约束的联合分布,生成符合人类时间动态的注视轨迹,在标准数据集上达到最先进性能。

中文摘要 AI 辅助

自我中心注视预测支持许多下游应用,但由于人类注视本质上是随机的,该任务仍然具有挑战性。这种随机性受到注视与扫视之间交替的结构化时间动态、来自任务的自上而下影响以及自下而上的视觉显著性的约束。基于这一观察,我们引入了GazeFlow,一个将注视直接建模为以自上而下和自下而上信息为条件的联合时间注视位置分布的框架。具体而言,GazeFlow使用条件流匹配(CFM):一个学习到的速度场迭代地将高斯噪声样本转化为从该联合分布中抽取的合理注视轨迹。速度场以视频编码器提取的自下而上视觉特征为条件,并通过全局查询这些特征获得自上而下的任务信息。在标准数据集上,GazeFlow在逐帧指标上达到了最先进的性能,且生成的轨迹与人类注视的时间动态更好地对齐。

英文摘要

Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.

发表机构

  • University of Rochester(罗切斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑