发表机构
Indian Institute of Technology Kharagpur; Indian Institute of Science Bangalore; École de technologie supérieure Montreal(印度理工学院卡哈拉格普尔分校; 印度科学理工学院班加罗尔分校; 蒙特利尔高等技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出VideoMSN框架,利用图像视觉变换器通过掩码孪生损失高效学习视频时空表示,在多个基准上达到最先进性能,并大幅减少预训练周期。
AI 中文摘要
我们提出了VideoMSN,一种掩码孪生网络框架,用于视频中高效的自监督时空表示学习。我们不依赖重型3D架构或基于重建的自编码器来利用未标记数据进行学习,而是通过将视频表示为超级图像(即由从视频中采样的帧组成的网格)来重新利用标准图像视觉变换器。从每个超级图像中,我们构建两个视图:一个带有空间补丁掩码,另一个带有时间帧掩码,确保帧之间没有信息泄漏。一个共享的视觉变换器(ViT)编码器使用掩码孪生损失对齐它们的嵌入,无需重建即可捕获运动和外观线索。我们的无解码器公式利用图像基础模型实现高效的视频表示学习。从预训练的DINO-v3和DeiT-v3图像编码器开始,VideoMSN在Kinetics-400、UCF101和HMDB51上实现了最先进的性能,同时与先前的视频自监督学习方法相比,所需的视频预训练周期分别减少了高达32倍和160倍。我们提出的方法在低样本分类中也表现出强大的性能,证实了在标签稀缺场景下所学表示的可迁移性。项目页面:此https URL。
英文摘要
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to $32\times$ fewer and $160\times$ fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.
CommentsAccepted in BMVC 2026