我有一条流:让自监督学习在连续视频上发挥作用
I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
- Faculty of Electrical Engineering and Computing, University of Zagreb(萨格勒布大学电气工程与计算学院)
- Fundamental AI Lab, University of Technology Nuremberg(纽伦堡工业大学基础人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出StreamMAE方法,通过流感知正则化和运动偏置裁剪解决连续视频流自监督学习中高批次内相似性的挑战,在95小时视频数据上达到与独立同分布MAE相当的性能。
AI中文摘要:
自监督学习从婴儿视觉发展中汲取灵感,然而标准的训练流程与之几乎没有相似之处:图像被独立采样,并在各轮训练中全局打乱。我们研究从连续视频流中进行自监督学习,其中帧按时间顺序使用严格的滑动窗口批次进行消费,不进行全局重排或多轮回放。为此,我们构建了WT++,一个95小时的城市步行游览视频数据集,用于流式预训练。结合一套全面的评估套件,我们发现基于对比和蒸馏的方法在此设置中表现不佳,而MAE更为稳健,但仍不及标准的独立同分布(i.i.d.)预训练。我们发现,由跨连续批次的滑动窗口消费导致的高批次间相似性并不能解释这一差距。主要挑战在于高批次内相似性,即每个批次内的帧几乎是重复的。为缓解这一问题,我们提出StreamMAE,它保留了MAE的核心重建目标,同时通过流感知正则化和运动偏置裁剪选择来调整输入流水线。StreamMAE优于流式基线,与在相同视频数据上训练的i.i.d. MAE相当,与ImageNet预训练的MAE保持竞争力,并且随着预训练流从12小时增长到95小时,性能呈正向扩展。
英文摘要:
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.