arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知道何时信任先验:用于视频注视预测的可靠性门控线索融合

Knowing When to Trust a Prior: Reliability-Gated Cue Fusion for Video Gaze Prediction

Lichen Zhu, Yueqian Lin, Yiheng Wang, Hai "Helen" Li, Yiran Chen

arXiv 2610.08663首次发表:更新:

发表机构

Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频注视预测,提出FocusGate门控融合框架,通过逐帧门控选择可信先验,显著提升多种监督预测器性能,仅增加1%延迟。

AI 中文摘要

视频注视预测由注视训练模型主导,然而无注视先验携带了这些模型未吸收的信号,前提是知道何时信任它们。我们提出FocusGate,一种无注视先验的门控集成,其成员可以弃权(不执行)。逐帧门控读取散焦图的三个形状统计量,并选择估计器在平均上高于偶然水平的帧,因此被拒绝的帧精确地降为基础,而中位数秩归一化使全零先验在零参数下弃权(不执行)。门控融合在电影、体育和网络视频上显著为正,而无条件融合在体育上有害,在网络视频上无效。添加到四个监督预测器(其中包括NTIRE 2026冠军)后,FocusGate在洗牌AUC上改善了所有十六个模型-领域单元,其中十五个显著,一个领域预先注册并评分一次,同时仅增加冠军延迟的1%。单独使用时,在电影上以16帧因果均值在洗牌AUC上超越TASED-Net和UNISAL。

英文摘要

Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the frames on which the estimator is above chance on average, so rejected frames reduce to the base exactly, while midrank normalisation lets an all-zero prior abstain at zero parameters. Gated fusion is significantly positive on film, sports and web video, whereas unconditional fusion is harmful on sports and null on web. Added to four supervised predictors, the NTIRE 2026 champion among them, FocusGate improves all sixteen model-domain cells in shuffled AUC, fifteen significantly, one domain pre-registered and scored once, while adding only 1% to the champion's latency. Alone, it surpasses TASED-Net and UNISAL in shuffled AUC on film with a 16-frame causal mean.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑