arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推进视频显著性研究:一个运动至关重要的新数据集与基准

Moving Forward with Video Saliency: A New Dataset and Benchmark where Motion Matters

Susmit Agrawal, Rebecca Wanner, Juliane Verwiebe, Matthias Tangemann, Matthias Bethge, Matthias Kümmerer

arXiv 2610.03276首次发表:更新:

发表机构

Tübingen AI Center, University of Tübingen; IMPRS-IS; Optocycle; University of Toronto; Vector Institute(图宾根大学图宾根人工智能中心; 国际智能系统马克斯·普朗克研究学院; Optocycle; 多伦多大学; 向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频显著性基准动态性不足的问题,提出高动态数据集SalTempto,验证时序模型能捕获更多时序信息,但仍有提升空间。

AI 中文摘要

视频显著性预测因额外的时序维度,本质上比静态图像显著性建模更难。视频显著性基准建立在这样的前提上:预测视频中的注视点需要利用分布在帧间的时序活动。先前的工作对此提出质疑,表明静态基线在流行的视频显著性数据集LEDOV上恢复了相当大比例的可解释注视信息,并且视频显著性模型在与该静态基线相同的地方失败。我们验证了这一诊断仍然成立:在比原始分析更强大的金标准和更新的近期架构面板下,面板中最强的时序架构并未显著优于微调后的静态基线。然而,目前尚不清楚这一边际增益反映的是当前时序架构的局限性,还是基准本身缺乏时序模式。我们引入了SalTempto,一个更具动态性的视频显著性基准:包含来自HACS-Segments数据集的224个高动态内容片段,每个片段包含一个事件及其前奏和后续,注视记录来自最多16名受试者,并提供训练分割以适应预训练模型。在SalTempto上,静态基线仅恢复了中心偏置之上可解释空间的约13%,而在LEDOV上则超过一半。微调后的时序架构相比静态基线在性能上有显著提升,表明它确实捕获了更多有意义的时序信息,而LEDOV无法衡量这一点。然而,即便是这一最先进的模型仍未能解释SalTempto近一半的可解释空间,表明视频显著性建模仍有改进余地。对SalTempto的检验也使我们能够描述人类被模型忽略的倾向。SalTempto链接:此https URL。

英文摘要

Video saliency prediction is inherently harder to model than static image saliency due to the additional temporal dimension. Video saliency benchmarks rest on the premise that predicting gaze on video requires utilizing temporal activity distributed across frames. Prior work has challenged this, showing that static baselines recover a significant fraction of the explainable gaze information on LEDOV, a popular video saliency dataset, and that video saliency models fail in the same places as this static baseline. We verify that this diagnosis still stands: under a more capable gold standard than the original analysis, and an updated panel of recent architectures, the strongest temporal architecture in the panel still does not substantially improve over a fine-tuned static baseline. However, it remains unclear whether the marginal gain reflects limitations of current temporal architectures or a lack of temporal patterns in the benchmark itself. We introduce SalTempto, a video saliency benchmark with greater dynamism: 224 clips of highly dynamic content, sourced from the HACS-Segments dataset so that each clip contains an event together with its lead-up and aftermath, with gaze recordings from up to 16 subjects and a training split for adapting pretrained models. On SalTempto, the static baseline recovers only about 13\% of the headroom above the centerbias, against more than half on LEDOV. A fine-tuned temporal architecture shows a substantial gain in performance over the static baseline, indicating that it does capture meaningfully more temporal information, which LEDOV fails to measure. Yet, even this SoTA model still leaves nearly half of SalTempto's headroom unexplained, indicating room for improvement in video saliency modelling. Examination of SalTempto also lets us describe human tendencies that models miss. SalTempto link: https://huggingface.co/datasets/bethgelab/video_saliency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑