arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2407.06871cs.CV

重新思考图像到视频的适配:一种以物体为中心的视角

Rethinking Image-to-Video Adaptation: An Object-centric Perspective

  • The Chinese University of Hong Kong(香港中文大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Rui Qian, Shuangrui Ding, Dahua Lin

更新

AI总结:

该研究从以物体为中心的视角提出高效图像到视频适配策略,通过槽注意力提炼物体token、物体-时间交互层建模时序变化,在动作识别上以更少参数达SOTA,零样本视频分割表现优异。

AI中文摘要:

图像到视频适配旨在高效地将图像模型适配应用于视频领域。许多图像到视频适配范式并未微调整个图像骨干网络,而是在空间模块之上使用轻量级适配器进行时序建模。然而,这些尝试在效率和可解释性方面存在局限。本文中,我们从以物体为中心的视角提出了一种新颖且高效的图像到视频适配策略。受人类感知将物体识别为视频理解关键组成部分的启发,我们将物体发现的代理任务整合到图像到视频的迁移学习中。具体而言,我们采用带有可学习查询的slot attention(槽注意力),将每帧提炼为一组紧凑的物体token(令牌)。这些以物体为中心的token随后通过物体-时间交互层处理,以建模物体随时间的状态变化。结合两种新颖的物体级损失,我们证明了仅在压缩的以物体为中心的表示上执行高效时序推理以完成视频下游任务的可行性。我们的方法在动作识别基准上取得了最优性能,且可调参数更少,仅为全微调模型的5%、高效调参方法的50%。此外,我们的模型在无需进一步重新训练或物体标注的零样本视频物体分割任务中表现优异,证明了以物体为中心的视频理解的有效性。

英文摘要:

Image-to-video adaptation seeks to efficiently adapt image models for use in the video domain. Instead of finetuning the entire image backbone, many image-to-video adaptation paradigms use lightweight adapters for temporal modeling on top of the spatial module. However, these attempts are subject to limitations in efficiency and interpretability. In this paper, we propose a novel and efficient image-to-video adaptation strategy from the object-centric perspective. Inspired by human perception, which identifies objects as key components for video understanding, we integrate a proxy task of object discovery into image-to-video transfer learning. Specifically, we adopt slot attention with learnable queries to distill each frame into a compact set of object tokens. These object-centric tokens are then processed through object-time interaction layers to model object state changes across time. Integrated with two novel object-level losses, we demonstrate the feasibility of performing efficient temporal reasoning solely on the compressed object-centric representations for video downstream tasks. Our method achieves state-of-the-art performance with fewer tunable parameters, only 5\% of fully finetuned models and 50\% of efficient tuning methods, on action recognition benchmarks. In addition, our model performs favorably in zero-shot video object segmentation without further retraining or object annotations, proving the effectiveness of object-centric video understanding.

补充信息

↑