arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LeVJEPA:无需启发式方法的高效可扩展视频预训练

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner

arXiv 2608.27395首次发表:更新:

发表机构

German Cancer Research Center; German Cancer Consortium; Goethe University Frankfurt; Mila; Université de Montréal; Brown University; Courant Institute, New York University; Advanced Machine Intelligence (AMI Labs)(德国癌症研究中心; 德国癌症联盟; 法兰克福大学; 米拉研究所; 蒙特利尔大学; 布朗大学; 纽约大学柯朗研究所; 高级机器智能实验室(AMI Labs))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LeVJEPA是首个基于LeJEPA无崩溃目标训练的视频编码器,无需启发式方法,预训练计算量大幅降低,在ImageNet-1K等基准上性能优于现有方法,可作为通用视觉预训练的更优基础。

AI 中文摘要

视频承载着物理世界的时间结构,但从视频中学习表示的计算成本一直很高:主流自监督方法要么通过架构不对称性防止表示崩溃,结合指数移动平均目标编码器、停止梯度和容量受限的预测器,要么通过在像素空间中重建掩码内容来规避崩溃。我们引入LeVJEPA,这是首个在LeJEPA无崩溃目标下训练的视频编码器,它无需上述两者。单个编码器通过对视频片段的全局和局部视图的不变性损失进行训练,该损失由SIGReg正则化,SIGReg具有可证明的保证可排除崩溃。该架构简化为一个编码器和一个投影器,目标简化为单个超参数。这种形式具备两个特性:首先,预训练成本由编码器观察到的令牌数量决定;均匀随机令牌丢弃使该数量变小,同时提高下游准确率。在相同数据的匹配轮次下,LeVJEPA在ViT-S/B/L上与V-JEPA 2相当或超越,且预训练计算量减少5.6至20.8倍;在匹配总FLOPs时,它在ImageNet-1K上比最强的视频基准高出7.6个点,同时在以运动为中心的基准上保持竞争力。其次,由于不需要分支之间的不对称性,编码器可以使用块因果注意力进行训练,且无明显准确率损失:时间顺序成为编码器本身的属性。与在相同视频帧上训练的计算匹配的DINOv2相比,LeVJEPA在以外观为中心的评估中接近图像预训练编码器,同时在以运动为中心的准确率上几乎翻倍。这些结果表明,一旦消除计算开销,视频就成为通用视觉预训练的可行且在多个方面更优的基础。

英文摘要

Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymmetries, coupling an exponential-moving-average target encoder, a stop-gradient, and a capacity-limited predictor, or circumvent it by reconstructing masked content in pixel space. We introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective, which dispenses with both. A single encoder is trained with an invariance loss over global and local views of a clip, regularized by SIGReg, which excludes collapse with a provable guarantee. The architecture reduces to an encoder and a projector, and the objective to a single hyperparameter. This formulation admits two properties. First, the cost of pretraining is governed by the number of tokens the encoder observes; uniform random token dropping renders this number small while simultaneously improving downstream accuracy. At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L at 5.6 to 20.8x less pretraining compute, and at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Second, since no asymmetry between branches is required, the encoder can be trained with block-causal attention at no measurable accuracy cost: temporal ordering becomes a property of the encoder itself. Against a compute-matched DINOv2 trained on frames of the same videos, LeVJEPA approaches the image-pretrained encoder on appearance-centric evaluation while nearly doubling its motion-centric accuracy. These results indicate that, once its computational overhead is removed, video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑