arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25465cs.CV

DensFiLM:用于人群场景的密度条件视频显著性

DensFiLM: Density-Conditioned Video Saliency for Crowd Scenes

Anis Ur Rahman

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对人群场景视频显著性模型单一注视策略问题,提出DensFiLM模型,在Video Swin Transformer瓶颈处插入轻量级层,利用密度嵌入生成参数重建显著性,增加少量参数,性能提升,表明该方法比增模型容量更有效。

中文摘要 AI 辅助

视频显著性模型通常在人群场景中采用单一的注视策略,尽管注意力会随着人群密度发生系统性变化。稀疏场景鼓励跟踪个体,而密集场景则将注意力转向集体运动和场景级地标。我们引入了DensFiLM,一种密度条件视频显著性模型,它在Video Swin Transformer的瓶颈处插入了一个轻量级的逐特征线性调制层。一个学习到的密度嵌入产生通道级的缩放和偏移参数,使解码器能够从为每个密度区域选择的特征中重建显著性。该模块仅增加约10万个参数,并且可以使用CrowdFix密度标签或模型自己的密度预测。在CrowdFix上,DensFiLM在四个种子上实现了平均NSS为1.434和CC为0.517,分别比ACLNet提高了14.7%和14.9%,而预测密度条件与神谕标签性能相匹配。消融实验表明,在这种设置下,显式的RAFT光流以及更大的时间和社会力扩展并没有进一步的改进。在中心先验减法诊断中,密度条件相比于无条件主干在NSS上有0.462的增益,而在标准评估下为0.124。这些结果表明,轻量级瓶颈条件比增加模型容量为人群视频显著性提供了更有效的归纳偏差。我们的代码可在这个https URL获取。

英文摘要

Video saliency models typically apply a single fixation strategy across crowd scenes, despite systematic changes in attention with crowd density. Sparse scenes encourage tracking individuals, whereas dense scenes shift attention toward collective motion and scene-level landmarks. We introduce DensFiLM, a density-conditioned video saliency model that inserts a lightweight Feature-wise Linear Modulation layer at the bottleneck of a Video Swin Transformer. A learned density embedding produces channel-wise scale and shift parameters, allowing the decoder to reconstruct saliency from features selected for each density regime. The module adds only ~100K parameters and can use either CrowdFix density labels or the model's own density prediction. On CrowdFix, DensFiLM achieves mean NSS 1.434 and CC 0.517 over four seeds, improving over ACLNet by 14.7% and 14.9%, respectively, while predicted-density conditioning matches oracle-label performance. Ablations show that explicit RAFT optical flow and larger temporal and social-force extensions provide no further improvement in this setting. In a centre-prior-subtraction diagnostic, density conditioning yields an NSS gain of 0.462 over the unconditioned backbone, compared with 0.124 under standard evaluation. These results show that lightweight bottleneck conditioning provides a more effective inductive bias than increasing model capacity for crowd-video saliency. Our code is available at https://github.com/aniskhan25/crowdfix-saliency.

发表机构

  • CSC - IT Center for Science Ltd.(CSC科学信息技术中心有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑