arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video-STLayout 预训练

Video-STLayout Pre-training

Akash Abdu Jyothi, Greg Mori

arXiv 2609.24031首次发表:更新:

发表机构

Simon Fraser University(西蒙弗雷泽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Video-STLayout预训练方法,利用物体检测器获取边界框时空布局,通过对比损失对齐视频与布局特征,提升复杂场景活动识别性能。

AI 中文摘要

近年来,预训练已成为学习有效视频表示的基础,能够实现向下游任务的强迁移。一种流行的预训练框架是将视频编码器的特征与另一模态(例如语言或音频)的特征进行对齐。我们提出了 Video-STLayout 预训练,这是一种新颖的策略,通过物体边界框的时空布局来获取丰富的视频表示。物体布局可以很容易地通过在视频帧上应用现成的物体检测器获得。我们的方法使用对比损失将视频特征与来自训练好的布局编码器的布局特征进行对齐。我们展示了该方法在复杂场景中的活动识别任务上的有效性。

英文摘要

In recent years, pre-training has become fundamental to learning effective video representations, enabling strong transfer to downstream tasks. A popular framework in pre-training involves aligning features of a video encoder with that of another modality, for example, language or audio. We introduce Video-STLayout pre-training, a novel strategy for obtaining rich video representations informed by spatio-temporal layout of object bounding boxes. Object layouts can easily be obtained by applying an off-the-shelf object detector on the video frames. Our method uses a contrastive loss to align video features with the layout features from a trained layout encoder. We show the effectiveness of our approach in the task of activity recognition in complex scenes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑