arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35288cs.LGcs.CV

$λ$-JEPA:用于自监督学习的谱抗坍缩正则化

$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning

Berker Demirel, Clémentine Dominé, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello

首次发表
浏览论文内容

中文总结 AI 辅助

提出SACReg谱抗坍缩正则化器,基于λ-平衡分析,应用于JEPA得到λ-JEPA,在图像和视频自监督学习中提升表征秩和下游迁移性能。

中文摘要 AI 辅助

联合嵌入自监督学习通常将跨增强视图的不变性目标与防止表征坍缩的额外机制相结合。这些目标通常在投影头之后应用,而下游任务使用投影器之前的骨干网络表征。我们发现这种不匹配并不一定能防止骨干网络中的维度坍缩,骨干网络可能保持较低的有效秩,并可能限制下游迁移性能。为了解决这个问题,我们引入了SACReg,一种谱抗坍缩正则化器,其动机来自对$\lambda$-平衡的分析,该分析捕捉了各层权重矩阵的相对尺度。在两层线性网络中,我们证明(i)$\lambda$-平衡防止坍缩,以及(ii)我们应用于骨干网络的正则化器诱导$\lambda$-平衡。在非线性情况下,该正则化器同样导致抗坍缩,并且在ImageNet100上的现实架构中,它经验性地增加了表征的秩。我们将SACReg应用于JEPA,并提出了$\lambda$-JEPA,它在ImageNet-1k分类和八个下游图像数据集的平均线性探针迁移性能上优于LeJEPA和VISReg。在视频自监督学习方面,$\lambda$-JEPA在Something-Something-v2和Kinetics-400基准上优于LeVJEPA和V-JEPA 2。代码可在以下网址获取:https://this URL。

英文摘要

Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of $λ$-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) $λ$-balance prevents collapse, and (ii) our regularizer applied to the backbone induces $λ$-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations' ranks. We apply SACReg to JEPA and propose $λ$-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, $λ$-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.

发表机构

  • Institute of Science and Technology Austria (ISTA)(奥地利科学技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑